Skip to content
Preprint

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

V-Link is proposed, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer and injects them into Action DiT through asymmetric pathways.

Abstract

Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.

View source

Similar papers

Preprint Aug 2026

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

AtVLA, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate, is introduced, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate.

Jin Cui, Yanbin Hu, Xinyue Long et al. · 0 citations
Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the range of scene poses that the demonstrations cover. We propose SAVLA, an end-to-end symmetry-aware VLA model for robust and data-efficient policy learning. Our approach keeps the pretrained vision-language backbone entirely frozen while combining it with an equivariant flow-matching action head and a learned canonicalizer. The head decomposes its state, action, and conditioning inputs into invariant and equivariant channels, and preserves this typing throughout all of its layers. The canonicalizer transforms oblique-view images into a canonical frame and rotates the geometric conditions consistently. We evaluate our model on LIBERO. Compared with the GR00T N1.5 baseline, SAVLA improves the success rate averaged over all four LIBERO suites by 5.1 points and increases the mean success rate under rotation on LIBERO-Goal from 41.5% to 90.4%.

Jun-Le Li, Weixian Waylon Li, Fuxiang Wu et al. · 0 citations
Preprint Aug 2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Space Tokens is introduced, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules, and demonstrates that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism for integrating geometric reasoning into large vision-language models.

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian et al. · 0 citations
Preprint Aug 2026

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

It is shown that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering, while approaching feature fusion methods with considerably fewer added parameters and lower latency.

K. T. Nguyen, Hanbo Shim, Jinwoo Kim et al. · 0 citations
Preprint Aug 2026

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.

Xing-Yu Ding, Yuzhong Zhao, Chunming Zhao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies.

Davood Soleymanzadeh, Kai-Di Zhang, Zhi-Yuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.