Skip to content
Preprint

Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

SALT is introduced, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents to substantially improve language-conditioned control.

Abstract

Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action space, where numerical proximity need not reflect linguistically meaningful distinctions. On BridgeV2, we show that action trajectories contain verb-grounding information beyond visual state changes, and that reconstruction-only discrete tokenization systematically erodes this information. To address this problem, we introduce SALT, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents. Policies trained with SALT achieve 71.9% average success in SimplerEnv, compared with 42.7% for a reconstruction-only VQ-VAE tokenizer and 31.2% for FAST. SALT also develops verb-specialized codes while maintaining reconstruction fidelity. These results show that robot action trajectories provide a source of language grounding and that preserving this structure in action representations can substantially improve language-conditioned control.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet sma...

Shijie Lian, Bin Yu, Zhaolong Shen et al. · 1 citation
Preprint Aug 2026

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

Sangoh Lee, Sang-Woo Mo, Wook-Shin Han · 0 citations
Preprint Aug 2026

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.

Xing-Yu Ding, Yu-Zhong Zhao, Chun-Ming Zhao et al. · 0 citations
Preprint Aug 2026

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

This work demonstrates that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm and introduces ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording.

Zhao-Kai Yin, Zhi-Peng Zhang · 0 citations
#artificial intelligence Preprint Aug 2026

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA polic...

Davood Soleymanzadeh, Kai-Di Zhang, Zhi-Yuan Zhang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.