Preprint
Aug 2026
How Should Vision-Language-Action Models Use Proprioceptive State?
Five representative interfaces are implemented -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluated on 45 atomic tasks spanning three task families plus 20 composite tasks.
Yiren Zhao, Ziyang Chen, Ziyang Rao et al.
· 0 citations