Structured-VLA: An Explicit Structural Interface for Vision-Language-Action Control
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but existing approaches often condition control on entangled multimodal features that do not explicitly preserve task-relevant structure during action generation. This limitation is especially pronounced in long-hori...