Skip to content
Conference

RoVLA: Action-Conditioned Temporal Grounding and Structured Conditioning for Vision–Language–Action Policies

Aug 2026 · 2026 IEEE 22nd International Conference on Automation Science and Engineering (CASE) · pp. 2219-2226 · 0 citations · 25 references

Abstract

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, yet pretrained policies do not always expose the information needed for precise action generation. In particular, effective control can benefit from (i) temporal context that reflects how recent actions have shaped the current scene, and (ii) perceptual cues organized according to the control roles implied by the task. We introduce RoVLA, a GR00T-N1-based enhancement that addresses these two issues with complementary modules: (i) an Action-Conditioned Encoder (ACE) that uses historical actions as temporal priors for action-conditioned grounding over state transitions, and (ii) a Control Injection (CI) module that injects triadic structured perceptual features into the pretrained policy head through a zero-initialized pathway for stable multimodal fusion. Experiments on the LIBERO benchmark show that RoVLA improves average success rate from 92.0% to 95.7% over its GR00T-N1 backbone, with especially clear gains on longhorizon and spatially demanding tasks. Real-world experiments on the SO-101 robot further show a 6.8-point improvement over GR00T-N1. Together, these results indicate that augmenting a fixed pretrained VLA backbone with explicit temporal and structured perceptual intermediate information provides a practical and effective way to improve robotic manipulation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.