V-Link is proposed, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer and injects them into Action DiT through asymmetric pathways.
Abstract
Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.
AtVLA, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate, is introduced, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate.
Jin Cui, Yanbin Hu, Xinyue Long et al.· 0 citations
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the range of scene poses that the demonstrations cover. We propose SAVLA, an end-to-end symmetry-aware VLA model for robust and data-efficient policy learning. Our approach keeps the pretrained vision-language backbone entirely frozen while combining it with an equivariant flow-matching action head and a learned canonicalizer. The head decomposes its state, action, and conditioning inputs into invariant and equivariant channels, and preserves this typing throughout all of its layers. The canonicalizer transforms oblique-view images into a canonical frame and rotates the geometric conditions consistently. We evaluate our model on LIBERO. Compared with the GR00T N1.5 baseline, SAVLA improves the success rate averaged over all four LIBERO suites by 5.1 points and increases the mean success rate under rotation on LIBERO-Goal from 41.5% to 90.4%.
Space Tokens is introduced, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules, and demonstrates that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism for integrating geometric reasoning into large vision-language models.
Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian et al.· 0 citations
It is shown that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering, while approaching feature fusion methods with considerably fewer added parameters and lower latency.
K. T. Nguyen, Hanbo Shim, Jinwoo Kim et al.· 0 citations
This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.
Xing-Yu Ding, Yuzhong Zhao, Chunming Zhao et al.· 0 citations
PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies.
Davood Soleymanzadeh, Kai-Di Zhang, Zhi-Yuan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.