World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...
Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al.· 0 citations
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdeterm...
Mengao Zhao, Ziang Li, Chaodong Huang et al.· 0 citations
EmbodiedGen V2 is established as scalable simulation infrastructure for training, evaluating, and deploying embodied policies, and is established as a generative, editable, and reusable simulation pipeline.
Xin-Jie Wang, Liu Liu, Taojun Ding et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.