Dynamic Sign Language Recognition: A Hybrid Approach Combining MediaPipe and LSTM
Keywords:
American Sign Language, Mediapipe Holistics, Deep-Learning, Sign Language Recognition, Long Short-Term Memory (LSTM)Abstract
The inability of the general population to understand sign language creates serious communication barriers for hearing-impaired individuals in critical situations such as healthcare emergencies, educational settings, and workplace environments. Addressing this gap requires effective and accurate automated recognition systems. This proposed research presents a real-time American Sign Language (ASL) recognition system for 24 dynamic signs that integrates the MediaPipe frame work with Long Short-Term Memory (LSTM) network. To enhance performance while reducing computational complexity, only the most relevant features are extracted from self-recorded dynamic sign videos: coordinates of 65 key hand and body landmarks, complemented by 41 engineered angle features between joint connections and 26 engineered distance features between specific landmark pairs, yielding a compact 285-feature representation per frame. As a result, LSTM network handles spatiotemporal sequence modeling across 25-frame sequences, effectively capturing the dynamic nature of sign language gestures. The proposed system achieves 98% test accuracy, with precision, recall, and F1-score of 98% across 720 test samples. It also successfully interprets all 24 dynamic signs in real-time testing scenarios using a standard webcam, including visually similar sign pairs such as “Good”/“Bad” and “Mother”/“Not”, demonstrating its practical applicability for assistive communication technologies. The research contribution lies in the systematic integration of MediaPipe Holistic with strategic feature engineering (combining landmark coordinates with computed angles and distances) and LSTM modeling to achieve efficient real-time dynamic sign recognition with reduced computational and data requirements.
References
[1] I. A. Adeyanju, O. O. Bello, and M. A. Adegboye, “Machine learning methods for sign language recognition: A critical review and analysis,” Intelligent Systems with Applications, vol. 12, pp. 1–36, 2021.
[2] S. A. E. El-Din and M. A. Abd El-Ghany, “Sign language interpreter system: An alternative system for machine learning,” in 2020 2nd Novel Intelligent and Leading Emerging Sciences Conference (NILES). Giza, Egypt: IEEE, Oct. 24–26, 2020, pp. 332–337.
[3] Y. Farhan, A. Ait Madi, A. Ryahi, and F. Derwich, “American sign language: Detection and automatic text generation,” in 2022 2nd International Conference on Innovative Research in Applied Science, Engineering and Technology (IRASET). Meknes, Morocco: IEEE, March 3–4, 2022, pp. 1–6.
[4] Y. Farhan and A. Ait Madi, “Real-time dynamic sign recognition using MediaPipe,” in 2022 IEEE 3rd International Conference on Electronics, Control, Optimization and Computer Science (ICECOCS). Fez, Morocco: IEEE, Dec. 1–2, 2022, pp. 1–7.
[5] B. Singh and A. K. Dubey, “Deaf education and sign language: Strategies, challenges, and benefits,” Naveen International Journal of Multidisciplinary Sciences (NIJMS), vol. 1, no. 3, pp. 75–82, 2025.
[6] S. Stone, “FSU expert highlights sign language’s role for the deaf and hard-of-hearing community,” 2025. [Online]. Available: https://bit.ly/3UbifaF
[7] N. A. A. Boi-Dsane, “Being understood: How to expand sign language access for the deaf community,” 2024. [Online]. Available: https: //bit.ly/3Ue3iEJ
[8] A. A. Hosain, P. S. Santhalingam, P. Pathak, H. Rangwala, and J. Kosecka, “Hand pose guided 3D pooling for word-level sign language recognition,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision. Computer Vision Foundation, 2021, pp. 3429–3439.
[9] A. S. Al-Shamayleh, R. Ahmad, N. Jomhari, and M. A. M. Abushariah, “Automatic arabic sign language recognition: A review, taxonomy, open challenges, research roadmap and future directions,” Malaysian Journal of Computer Science, vol. 33, no. 4, pp. 306–343, 2020.
[10] S. Srivastava, S. Singh, Pooja, and S. Prakash, “Continuous sign language recognition system using deep learning with MediaPipe holistic,” Wireless Personal Communications, vol. 137, pp. 1455–1468, 2024.
[11] S. Alyami, H. Luqman, and M. Hammoudeh, “Isolated Arabic sign language recognition using a transformer-based model and landmark keypoints,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 1, pp. 1–19, 2024.
[12] A. Sultan, W. Makram, M. Kayed, and A. A. Ali, “Sign language identification and recognition: A comparative study,” Open Computer Science, vol. 12, no. 1, pp. 191–210, 2022.
[13] R. A. Kadhim and M. Khamees, “A real-time American sign language recognition system using convolutional neural network for real datasets,” TEM Journal, vol. 9, no. 3, pp. 937–943, 2020.
[14] A. Costa, “ASLScribe: Real-time American sign language alphabet image classification using MediaPipe hands and artificial neural networks,” 2019, final project in University of Georgia.
[15] J. Galka, M. Masior, M. Zaborski, and K. Barczewska, “Inertial motion sensing glove for sign language gesture acquisition and recognition,” IEEE Sensors Journal, vol. 16, no. 16, pp. 6310–6316, 2016.
[16] P. Kumar, H. Gauba, P. P. Roy, and D. P. Dogra, “Coupled HMM-based multi-sensor data fusion for sign language recognition,” Pattern Recognition Letters, vol. 86, pp. 1–8, 2017.
[17] F. Obaid, A. Babadi, and A. Yoosofan, “Hand gesture recognition in video sequences using deep convolutional and recurrent neural networks,” Applied Computer Systems, vol. 25, no. 1, pp. 57–61, 2020.
[18] M. U. Rehman et al., “Dynamic hand gesture recognition using 3D-CNN and LSTM networks,” Computers, Materials & Continua, vol. 70, no. 3, pp. 4675–4690, 2022.
[19] W. Zhang, J. Wang, and F. Lan, “Dynamic hand gesture recognition based on short-term sampling neural networks,” IEEE/CAA Journal of Automatica Sinica, vol. 8, no. 1, pp. 110–120, 2020.
[20] A. Kanodia, P. Singh, D. Rajesh, and G. Malathi, “Indian sign language using holistic pose detection,” Turkish Journal of Physiotherapy and Rehabilitation, vol. 32, no. 3, pp. 19 746–19 752, 2021.
[21] B. Duy Khuat, D. Thai Phung, H. Thi Thu Pham, A. Ngoc Bui, and S. Tung Ngo, “Vietnamese sign language detection using MediaPipe,” in Proceedings of the 2021 10th International Conference on Software and Computer Applications. Kuala Lumpur, Malaysia: Association for Computing Machinery, Feb. 23–26, 2021, pp. 162–165.
[22] A. Chaikaew, K. Somkuan, and T. Yuyen, “Thai sign language recognition: an application of deep neural network,” in 2021 Joint International Conference on Digital Arts, Media and Technology with ECTI Northern Section Conference on Electrical, Electronics, Computer and Telecommunication Engineering. Cha-am, Thailand: IEEE, March 3–6, 2021, pp. 128–131.
[23] M. G. Grif and Y. K. Kondratenko, “Development of a software module for recognizing the fingerspelling of the Russian sign language based on LSTM,” Journal of Physics: Conference Series, vol. 2032, no. 1, pp. 1–7, 2021.
[24] J. Shin, A. Matsuoka, M. A. M. Hasan, and A. Y. Srizon, “American sign language alphabet recognition by extracting feature from hand pose estimation,” Sensors, vol. 21, no. 17, pp. 1–19, 2021.
[25] L. Zheng, B. Liang, and A. Jiang, “Recent advances of deep learning for sign language recognition,” in 2017 International Conference on Digital Image Computing: Techniques and Applications (DICTA). Sydney, NSW, Australia: IEEE, Nov. 29–Dec. 01, 2017, pp. 1–7.
[26] N. B. Ibrahim, H. H. Zayed, and M. M. Selim, “Advances, challenges and opportunities in continuous sign language recognition,” Journal of Engineering and Applied Sciences, vol. 15, no. 5, pp. 1205–1227, 2020.
[27] C. Lugaresi et al., “MediaPipe: A framework for building perception pipelines,” 2019. [Online]. Available: https://arxiv.org/abs/1906.08172
[28] S. Bouktif, A. Fiaz, A. Ouni, and M. A. Serhani, “Multi-sequence LSTM-RNN deep learning and metaheuristics for electric load forecasting,” Energies, vol. 13, no. 2, pp. 1–21, 2020.
[29] A. Sherstinsky, “Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network,” Physica D: Nonlinear Phenomena, vol. 404, 2020.
[30] GitHub, “MediaPipe holistic,” 2024. [Online]. Available: https://github.com/google-ai-edge/mediapipe/blob/master/docs/solutions/holistic.md
[31] ——, “MediaPipe hands,” 2024. [Online]. Available: https://github.com/google-ai-edge/mediapipe/blob/master/docs/solutions/hands.md
[32] ——, “MediaPipe pose,” 2024. [Online]. Available: https://github.com/google-ai-edge/mediapipe/blob/master/docs/solutions/pose.md
[33] B. Guruprakash, N. Gurusamy, M. Ramnath, S. Sumathi, E. Mariappan, and T. Saravanan, “ISL sign language recognition using LSTMdriven deep learning model,” Journal of Electrical Systems, vol. 20, no. 3, pp. 1–9, 2024.
[34] Ridwang, A. A. Ilham, I. Nurtanio, and Syafaruddin, “Dynamic sign language recognition using mediapipe library and modified LSTM method,” International Journal on Advanced Science, Engineering & Information Technology, vol. 13, no. 6, pp. 2171–2180, 2023.
[35] M. Z. Uddin, C. Boletsis, and P. Rudshavn, “Realtime Norwegian sign language recognition using MediaPipe and LSTM,” Multimodal Technologies and Interaction, vol. 9, no. 3, pp. 1–15, 2025.
[36] G. Khartheesvar, M. Kumar, A. K. Yadav, and D. Yadav, “Automatic Indian sign language recognition using MediaPipe holistic and LSTM network,” Multimedia Tools and Applications, vol. 83, pp. 58 329–58 348, 2024.
[37] A. M. Buttar et al., “Deep learning in sign language recognition: A hybrid approach for the recognition of static and dynamic signs,” Mathematics, vol. 11, no. 17, pp. 1–20, 2023.
[38] A. Baihan, A. I. Alutaibi, M. Alshehri, and S. K. Sharma, “Sign language recognition using modified deep learning network and hybrid optimization: A Hybrid Optimizer (HO) based optimized CNNSa-LSTM approach,” Scientific Reports, pp. 1–22, 2024.
[39] E. Hassan, M. Y. Shams, T. Abd El-Hafeez, and M. Elseddik, “A novel model for expanding horizons in sign language recognition,” Scientific Reports, vol. 15, pp. 1–21, 2025.
[40] Y. Farhan and A. Ait Madi, “DynASL-24: Dynamic American Sign Language (ASL) recognition dataset,” 2026. [Online]. Available: https://zenodo.org/records/21462002
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Youssef Farhan, Zineb Haimer, Oumaima Lamaamar, Abdessalam Ait Madi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Â
USER RIGHTS
All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows: Creative Commons Attribution-Share Alike (CC BY-SA)

















