Aug 2026· International Journal of Electronics and Communication Engineering· 0 citations· 26 references
TL;DR
The results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.
Abstract
Continuous Sign Language Recognition (CSLR) has always stood quite difficult due to issues like coarticulation effects, availability of weak temporal annotations, and large variability among different signers, especially in a limited-resource scenario such as Indian Sign Language (ISL). Most of the current methods use end-to-end sequence modeling, which leaves out the explicit temporal structure and does not work well with weak supervision. Here, we propose a structured CSLR system that combines motion-guided segmentation, multimodal representation learning, and position-aware decoding aimed at overcoming these issues. More importantly, we present an MGPT-based temporal segmentation method that uses optical-flow-driven motion signals and Gaussian peak modeling to separate continuous signing sequences into consistent motion segments, which results in the reduction of transitional ambiguity. The spatial-temporal features are obtained with the help of a dual-stream architecture that integrates ResNet50-based visual representations and skeleton keypoint features, being then temporally modeled by a multi-layer LSTM network. To improve sequence-level consistency, we also introduce a Word Position Graph (WPG) for structured decoding along with Gaussian-weighted frame voting to highlight informative temporal regions and, at the same time, downplay noisy transitions. The approach we suggested was tested on the ISL-CSLRT dataset with weak sentence-level supervision. The experimental results show that our framework reaches 92% accuracy and a Word Error Rate (WER) of 0.07, greatly beating the baseline voting strategies. Statistical verifications, including multi-run evaluation and significance testing, have confirmed the robustness of the improvements. Also, comparing with representative CSLR methods has shown that the method of explicit temporal segmentation and position-aware decoding is very effective, especially when the dataset is scarce. Besides, the results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.
This work repurposes the CTC decoder to restrict representation-based scoring to decoder-aligned gloss regions, rather than exposing the acquisition function to the entire unfiltered video, and introduces RAIDAL, which achieves its strongest data-efficiency gains over competing baselines in large-vocabulary, budget-lim...
R. A. Diniz Augusto, Gabriel L. Oliveira, Erickson R. Nascimento· 0 citations
With the growing emphasis on accessibility-oriented technologies and inclusive intelligent systems, Continuous Sign Language Recognition (CSLR) has attracted increasing attention as a key technique for bridging communication between Deaf and hearing communities. However, existing methods still suffer from insufficient...
Ya-Han Yang, Rui Wang, Xiao-Fang Li et al.· International Conferences on...· 0 citations
SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evalu...
Jun-Yi Hu, Zhe-Wen He, Hao Huang et al.· 0 citations
Sign language is an essential communication system for hearing-impaired individuals, which mainly depends upon complex hand gestures and facial expressions. Automating Sign Language Recognition (SLR) from videos can enhance barrier-free communication, yet it remains challenging due to the subtle nature of signs in dive...
A. Babisha, G. Srikanth· Discover Computing· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.