Preprint
Aug 2026
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
SALT is introduced, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents to substantially improve language-conditioned control.
Wen-Jie Li, Yash Jangir, Ignacy Stępka et al.
· 0 citations