Automated Conversion of Arabic Sign Language to Text Based on the Hybrid CNN-LSTM Model to Enable Communication with the Deaf and Non-Speaking Individuals
Abstract
Communicating with the Deaf and non-speaking individuals and understanding sign language are complicated processes. With the continuous development in deep learning and automation of systems, however, solutions are available to overcome the difficulty of communicating with this segment of society and understanding sign language, especially in Arabic-speaking regions. This article proposes an automated sign language recognition system based on a hybrid convolutional neural network (CNN)–long short-term memory (LSTM) architecture, which represents two layers: CNN, which extracts spatial features from sign language motions, and LSTM, which records the temporal connections within gesture sequences. The model’s ability to provide smooth communication is demonstrated by evaluating its performance using a specially created Arabic Sign Language (ArSL) dataset, and then the model was tested and trained by using input videos recorded for three signers, each one of whom pronounced 50 different signs. We added three other signers to pronounce the days of the week in sign language, improving the accuracy of the pretraining. These videos served as added data to the ArSL dataset. The key point was extracted by using matplotlib, the key point values for training and testing were collected, and the data were preprocessed through feature extraction and batch normalisation. The signs were recognised in the Arabic language. The proposed model detects the sign language action and recognition according to the dataset training, which was achieved building and training by using the LSTM neural network and CNN. The accuracy rating, F-1 score, recall and precision of the proposed model for ArSL were 95%, 95%, 95% and 96%, respectively, with 126 epochs and batch normalisation equal to 30. Ultimately, the proposed model was able to recognise and detect ArSL accurately.