Skip to content
Preprint

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

This paper argues that what a VLA needs is not the ability to generate language, but the ability to consume grounded language, and introduces a framework that endows a VLA with language competence through in-context post-training and an agentic tool-use interface.

Abstract

Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expert action chunks conditioned on a static image and a fixed instruction. A natural remedy is to inject explicit reasoning through textual chain-of-thought (CoT). We show, both empirically and analytically, that free-form textual CoT degrades low-level control: the reasoning it produces is ungrounded, its latency breaks closed-loop timing, and, crucially, the reasoning and action tokens are optimized against conflicting objectives so that the policy learns to narrate rather than to act. We argue that what a VLA needs is not the ability to generate language, but the ability to consume grounded language. To this end we introduce \textbf{\ourmethod{}}, a framework that endows a VLA with language competence through (i) in-context post-training, in which perceptual evidence is injected as structured context and the model is supervised only on actions, and (ii) an agentic tool-use interface, in which the policy queries open-vocabulary detectors, monocular depth, and a vision--language model to actively acquire task-relevant information. Rather than emitting a single templated caption, our data engine produces diverse, paraphrased, and evidence-conditioned spatial descriptions, so that the policy learns to interpret language it has never seen verbatim. Across the RoboCasa-GR1, SimplerEnv, and LIBERO simulation benchmarks, together with 8 real-world robot manipulation tasks, our method consistently achieves SOTA results in both performance and efficiency when compared with CoT-based approaches under matched configurations.

View source

Similar papers

Preprint Aug 2026

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

Sangoh Lee, Sangwoo Mo, Wook-Shin Han · 0 citations
#artificial intelligence Preprint Sep 2026

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

Adding an auxiliary supervised contrastive (InfoNCE) loss during world-model warmup substantially improves sleep-phase clustering, and adding an auxiliary supervised contrastive loss during world-model warmup substantially improves sleep-phase clustering on LIBERO.

Riyaaz Shaik, Chandru Venkataraman · 0 citations
Preprint Aug 2026

How Should Vision-Language-Action Models Use Proprioceptive State?

Five representative interfaces are implemented -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluated on 45 atomic tasks spanning three task families plus 20 composite tasks.

Yiren Zhao, Ziyang Chen, Zi-Yang Rao et al. · 0 citations
Preprint Aug 2026

G0.5: One Autoregressive Stream for Robot Reasoning and Action

G0.5 is introduced, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective, which exceeds state-of-the-art models across 7 independent regimes.

Yicheng Liu, Zibin Dong, Baijun Ye et al. · 5 citations · ⚡1
Preprint Aug 2026

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

This work demonstrates that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm and introduces ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording.

Zhaokai Yin, Zhi-Peng Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.