Skip to content
#computer vision Preprint

TransSLR: A Lightweight Transformer for Sign Language Recognition

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

This work proposes TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head that achieves signer-independent generalization without relying on visual appearance on the CASL-W60 benchmark.

Abstract

Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we show that the common heuristic of fine-tuning high-resource models fails to close it. This failure stems from two compounding factors: the limited scale of available CASL data and the significant lexical and visual domain gap between CASL and large-scale corpora such as WLASL, which renders pre-trained representations largely uninformative. To address this, we propose TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves signer-independent generalization without relying on visual appearance. On the CASL-W60 benchmark, TransSLR establishes a new state-of-the-art accuracy of 80.39%, surpassing the prior best by +10.46%. Beyond accuracy, our encoder-only design significantly reduces computational overhead, making deployment feasible in resource-constrained environments. We conduct extensive experiments on the CASL-W60 benchmark, comparing against RGB-based and multimodal baselines, and demonstrate that TransSLR achieves state-of-the-art performance.

View source

Similar papers

#computer vision Preprint Sep 2026

Zero-Shot Cross-Lingual Recognition of Sign Language Handshapes

This work presents the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC), and leverages the decomposition of handshapes into five phonological features shared across both languages, to decode LSC handshapes from predicted features via a composite pho...

Marcel Granero-Moya, Carolina del Corral Farrarós, G. Haro et al. · 0 citations
2026

From text to sign language pose sequences: a Transformer-based approach for the automatic generation of gestures

A Transformer-based approach for the automatic generation of sign language pose sequences directly from text without relying on gloss-based intermediate representations is proposed, demonstrating the potential of attention-based architectures for sign language generation and highlighting their relevance for improving d...

M. H. H. Sezan, Pélagie Houngue · 0 citations
Conference Aug 2026

A Transformer-Based Deep Learning Framework for Real Time Sign Language Recognition and Multilingual Translation

Communication has been an important aspect in all the human lives. Especially for persons with disability to communicate linguistically, sign language comes to the rescue. Bridging the gap between the signers and non-signers is essential. To enable the signers to express in a language becomes essential while communicat...

Priyadarsini Plk, M. P, Priyadharshan M et al. · 0 citations
Conference Aug 2026

A Greedy Skeleton Retrieval Framework for Vietnamese Text-to-Sign Generation

Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...

D. Thanh, Thang Cap · 0 citations
Conference Aug 2026

VSL-HV8: A Vietnamese Sign Language Dataset for Sign Recognition and Translation Research

Vietnamese Sign Language (VSL) remains under-resourced for automatic sign language recognition and translation research. Existing VSL datasets primarily focus on isolated gloss-level recordings, while a large-scale dataset containing continuous sentence-level signing sequences is still unavailable. This limitation rest...

Hien Phuong Nhat Nguyen, Hoai Nhan Nguyen, Thi Diem Tran · 0 citations
Preprint Sep 2026

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.

Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.