Skip to content
Conference

VSL-HV8: A Vietnamese Sign Language Dataset for Sign Recognition and Translation Research

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 742-747 · 0 citations · 21 references

Abstract

Vietnamese Sign Language (VSL) remains under-resourced for automatic sign language recognition and translation research. Existing VSL datasets primarily focus on isolated gloss-level recordings, while a large-scale dataset containing continuous sentence-level signing sequences is still unavailable. This limitation restricts the development of models that can exploit temporal context across sign sequences and prevents standardized evaluation for continuous sign language recognition and translation in VSL. In this paper, we introduce VSL-HV8, a Vietnamese Sign Language dataset constructed from continuous sentence recordings with gloss-level segmentation annotations. The dataset consists of 4,160 continuous sentence videos collected from 8 signers and is further segmented into 15,792 isolated gloss clips covering 2,062 unique glosses. Unlike previous VSL datasets that only provide isolated sign samples, the sentence-level structure of VSL-HV8 preserves contextual relationships between consecutive signs, enabling isolated sign recognition while establishing a foundation for future continuous recognition and sign language translation research. To establish initial benchmarks, we construct three experimental subsets with vocabulary sizes of 50, 100, and 200 glosses and evaluate an I3D model pretrained on AUTSL. The proposed baseline achieves Top-1 accuracies of 75.99% on the 50-gloss subset and 61.48% on the 200gloss subset. We further provide a preliminary signer-independent (cross-signer) evaluation and discuss remaining limitations, including signer imbalance, long-tailed class distributions, and the lack of Vietnamese text translation annotations. VSL-HV8 is intended as an initial resource for advancing Vietnamese sign language recognition and translation research rather than a complete benchmark for robust cross-signer or end-to-end translation performance.

View source

Similar papers

Open access Aug 2026

A Multi-view Dataset for Vietnamese Word-Level Sign Language Recognition

This research introduces VSL400, a multi-view video dataset for isolated word-level recognition of Vietnamese Sign Language (VSL). The dataset contains 74,259 manually annotated video clips covering 400 glosses performed by 28 signers, including deaf student signers and additional trained VSL signers. Each signing in...

Trung Nguyen Quoc, Khoi Pham Dang, Viet Truong Duy et al. · 1 citation
#natural language process... Preprint Sep 2026

Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting

We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (\'ITM). We use \'ITM SignWiki, a dataset derived from a bilingual Icelandic--\'ITM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the f...

F. Ingimundarson, Gudny Bjork Thorvaldsdottir, Mathias Müller et al. · 0 citations
Conference Aug 2026

A Greedy Skeleton Retrieval Framework for Vietnamese Text-to-Sign Generation

Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...

D. Thanh, Thang Cap · 0 citations
#computer vision Preprint Sep 2026

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...

Zhaoyi An, Si-Han Tan, Youngbae Hwang et al. · 0 citations
Preprint Sep 2026

Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos

Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segme...

Bo-Wen Guo, Shi-Wei Gan, Ya-Feng Yin et al. · 0 citations
Preprint Sep 2026

SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

A prototype-structured sign embedding space is learned from continuous video annotated with signs, where each learnable prototype corresponds to a sign class, enabling the matching between dictionary exemplars and continuous sign instances.

Ryan Wong, Youngjoon Jang, Liliane Momeni et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.