Skip to content

VTaMo: Video-Text Alignment Model for Sign Language Translation

Jul 2026 · arXiv.org · Vol abs/2607.09126 · 0 citations · 47 references
Computer Science

TL;DR

VTaMo is presented, a framework that introduces explicit multi-granularity alignment at three levels: local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance.

Abstract

Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation supervision alone. We present VTaMo, a framework that introduces explicit multi-granularity alignment at three levels: (1) local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; (2) global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance; and (3) position-aligned contrastive learning for discriminative token-level representations. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL demonstrate consistent state-of-the-art performance, with ablations confirming the complementary contributions of each component. Code is available at https://github.com/junyi2005/vtamo.

View source

Similar papers

Aug 2026

Variational Sign Language Translation

Rui Zhao, Liang Zhang, Biao Fu et al. · 0 citations
Preprint Jul 2026

Attention-Steered Vision-Language Models for Sign Language Translation

Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased translators: poor spatial-temporal visual grounding. In particular, we find that standard n...

Meibo Hu, Guohao Sun, Annemarie D. Ross et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation

American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit se...

Hongyu Wu, Xu Wu, Tianhao Wu et al. · 0 citations
Jul 2026

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine r...

Hongbin Zhang, Jun-Hao Liu, Xue-Feng Bai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation that aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4.

Oline Ranum, Edward Fish, Simon Hadfield et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.