A Real-Time Deep Learning-Based Sign Language Recognition System for Words, Alphabets, and Numbers
Abstract
This paper presents a sign language recognition system based on deep learning and computer vision. It aims to support communication between deaf and hard-of-hearing individuals and the community. The proposed system translates hand gestures into textual output through real-time image and video processing. It supports three recognition modes, including number, alphabet, and word recognition. For alphabet and digit recognition, MobileNetV2 is adopted due to its lightweight architecture and suitability for real-time deployment. To improve recognition accuracy and robustness in practical scenarios, a custom-collected dataset is combined with publicly available American Sign Language (ASL) datasets. The resulting dataset covers static alphabet gestures (excluding the motion-based letters J and Z) and digit gestures corresponding to the numbers 0–9. For word recognition, the pretrained Inflated 3D ConvNet (I3D) model is adopted to process short video clips and learn both spatial and temporal gesture features. Additionally, OpenCV is employed for video processing and real