Structured-VLA: An Explicit Structural Interface for Vision-Language-Action Control
Abstract
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but existing approaches often condition control on entangled multimodal features that do not explicitly preserve task-relevant structure during action generation. This limitation is especially pronounced in long-horizon or relationally ambiguous tasks, where reliable control requires persistent grounding of object roles, spatial targets, and interaction intent. We present Structured-VLA, a framework that introduces a fixed-format structural interface between multimodal understanding and continuous control. A structure-aware VisionLanguage Model predicts predefined fields encoding semantic entities, geometric task anchors, agent-centric relations, and latent action codes. Their field-aligned hidden states are then reused as compact conditioning signals for a DiT-based Flow Matching policy, while dense visual features are preserved separately for local geometric refinement. Across LIBERO, Simpler-WidowX, and real-world SO101 evaluation, Structured-VLA consistently improves over a reproduced GR00T-N1.5 baseline, raising average success from 94.7% to 96.5% on LIBERO, from 39.8% to 45.1% on SimplerWidowX, and from 73.1% to 77.3% on SO101. The gains are particularly notable on long-horizon, relation-aware, and multistep manipulation tasks, while controlled ablations suggest that both the structural interface and the policy-side conditioning design contribute to the observed improvement. Overall, these results suggest that explicit structural conditioning is a promising way to improve downstream control performance and cross-domain robustness in pretrained VLA systems for robotic manipulation.