A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a c...
Chu-Yao Fu, Xiao-Wei Chi, Yu-Han Rui et al.· 0 citations
TrAct is proposed, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction, enabling more accurate world modeling and stronger robot generalization.
Zhi-Hang Cao, Howard Ji, Kevin Zhang et al.· 3 citations
This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...
Yifan Ye, Yankai Fu, Ya-hui Lv et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.