Skip to content

Multi-view Integration with View-specific Self-attention for Sign Language Recognition

2026 · Journal of Image and Graphics · 0 citations · 53 references

Abstract

— Sign Language Recognition (SLR) plays a pivotal role in mitigating communication barriers between the deaf community and broader society. Recently, the Video Transformer Network (VTN), an extension of the transformer architecture incorporating a multi-head self-attention mechanism, has demonstrated efficacy in video processing tasks and SLR. However, relying solely on Red Green Blue (RGB) frame information may render the model susceptible to redundant data and environmental factors such as lighting variations and complex backgrounds. Additionally, most existing SLR datasets provide only frontal viewpoints, which constrain the model ’ s generalizability, particularly in real-world settings where diverse perspectives are prevalent. In this study, we propose Video Transformer Network with 3 Graph Convolutional Network (VTN3GCN), a multi-view and multi-stream framework that integrates RGB data, skeleton coordinates, and pose flow from three distinct viewpoints: left, right, and center. This architecture enhances VTN by incorporating Graph Convolutional Networks (GCN) to learn skeletal frame features and employs an early fusion mechanism between the RGB and skeleton streams. Experiments conducted on the Multi-VSL200 dataset — a newly curated dataset for Vietnamese Sign Language (VSL) featuring three viewpoints per video demonstrate that the VTN3GCN framework achieves a top-1 accuracy of up to 92.92%, achieving state-of-the-art accuracy. Our code and data are available at: https://github.com/fossbk/MultiView-ISLR/tree/main/VTN3 GCN.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.