A Lightweight Hybrid 1D CNN and Bi GRU Architecture for Real Time Sign Language Recognition
Abstract
An efficient process for recognizing sign languages plays an important role in reducing the communication barrier between hearing-impaired people and others. Nevertheless, most of the currently available sign language recognition processes fail to produce highly accurate results within a minimum time period. In this paper, we perform an analysis of machine learning and deep learning approaches based on skeletal coordinates for the purpose of achieving a high level of precision and minimal latency. We implement the MediaPipe algorithm that detects 126 body landmarks to minimize the complexity of computations without using pixels. We consider random forest and KNN algorithms for the traditional approach alongside LSTM, GRU, and Transformers in terms of temporal deep learning algorithms. Besides that, we introduce our proposed solution with the implementation of a one-dimensional CNN and bidirectional gated recurrent unit. It has been demonstrated that our proposed method is capable of producing an accuracy of 99.31 percent in just 5.57 ms, while the best of traditional algorithms was able to reach an accuracy of 93.90 percent.