World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...
Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al.· 0 citations
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by...
Yue-Ting Zhu, Shao-Yu Chen, Yue-Hao Song et al.· 0 citations
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy tha...
Si-Xu Yan, Shi-Kang Wang, Bin-Hua Huang et al.· 1 citation
DreamWAM is introduced, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics, showing that robust world-action learning depends not only on predicting the future, but on representing it in a fo...
Shanglin Yuan, Weiheng Zhao, Xin Shi et al.· 4 citations
ABot-C0 is presented, a generalist motion-control system for quadruped robots that establishes three complementary behavior foundations: a scalable multi-source motion-data pipeline, robust policy learning across motion tracking, locomotion, and scene interaction, and a unified deployment stack for reliable real-world...