Skip to content
Conference

Structured-VLA: An Explicit Structural Interface for Vision-Language-Action Control

Aug 2026 · 2026 IEEE 22nd International Conference on Automation Science and Engineering (CASE) · pp. 2227-2234 · 0 citations · 19 references

Abstract

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but existing approaches often condition control on entangled multimodal features that do not explicitly preserve task-relevant structure during action generation. This limitation is especially pronounced in long-horizon or relationally ambiguous tasks, where reliable control requires persistent grounding of object roles, spatial targets, and interaction intent. We present Structured-VLA, a framework that introduces a fixed-format structural interface between multimodal understanding and continuous control. A structure-aware VisionLanguage Model predicts predefined fields encoding semantic entities, geometric task anchors, agent-centric relations, and latent action codes. Their field-aligned hidden states are then reused as compact conditioning signals for a DiT-based Flow Matching policy, while dense visual features are preserved separately for local geometric refinement. Across LIBERO, Simpler-WidowX, and real-world SO101 evaluation, Structured-VLA consistently improves over a reproduced GR00T-N1.5 baseline, raising average success from 94.7% to 96.5% on LIBERO, from 39.8% to 45.1% on SimplerWidowX, and from 73.1% to 77.3% on SO101. The gains are particularly notable on long-horizon, relation-aware, and multistep manipulation tasks, while controlled ablations suggest that both the structural interface and the policy-side conditioning design contribute to the observed improvement. Overall, these results suggest that explicit structural conditioning is a promising way to improve downstream control performance and cross-domain robustness in pretrained VLA systems for robotic manipulation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.