This work proposes TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head that achieves signer-independent generalization without relying on visual appearance on the CASL-W60 benchmark.
Abstract
Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we show that the common heuristic of fine-tuning high-resource models fails to close it. This failure stems from two compounding factors: the limited scale of available CASL data and the significant lexical and visual domain gap between CASL and large-scale corpora such as WLASL, which renders pre-trained representations largely uninformative. To address this, we propose TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves signer-independent generalization without relying on visual appearance. On the CASL-W60 benchmark, TransSLR establishes a new state-of-the-art accuracy of 80.39%, surpassing the prior best by +10.46%. Beyond accuracy, our encoder-only design significantly reduces computational overhead, making deployment feasible in resource-constrained environments. We conduct extensive experiments on the CASL-W60 benchmark, comparing against RGB-based and multimodal baselines, and demonstrate that TransSLR achieves state-of-the-art performance.
This work presents the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC), and leverages the decomposition of handshapes into five phonological features shared across both languages, to decode LSC handshapes from predicted features via a composite pho...
Marcel Granero-Moya, Carolina del Corral Farrarós, G. Haro et al.· 0 citations
A Transformer-based approach for the automatic generation of sign language pose sequences directly from text without relying on gloss-based intermediate representations is proposed, demonstrating the potential of attention-based architectures for sign language generation and highlighting their relevance for improving d...
M. H. H. Sezan, Pélagie Houngue· CITA-DW· 0 citations
Communication has been an important aspect in all the human lives. Especially for persons with disability to communicate linguistically, sign language comes to the rescue. Bridging the gap between the signers and non-signers is essential. To enable the signers to express in a language becomes essential while communicat...
Priyadarsini Plk, M. P, Priyadharshan M et al.· International Conference on...· 0 citations
Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...
D. Thanh, Thang Cap· International Conference on...· 0 citations
Vietnamese Sign Language (VSL) remains under-resourced for automatic sign language recognition and translation research. Existing VSL datasets primarily focus on isolated gloss-level recordings, while a large-scale dataset containing continuous sentence-level signing sequences is still unavailable. This limitation rest...
Hien Phuong Nhat Nguyen, Hoai Nhan Nguyen, Thi Diem Tran· International Conference on...· 0 citations
SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.
Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.