Whole-body teleoperation requires a humanoid robot to reproduce a human operator's behavior even when their terrains differ. This demands that the robot perceive local terrain and adapt its posture and contacts accordingly, rather than copy the operator's motion frame by frame. However, paired motion data linking the s...
Xiang-Yu Miao, Jun-Song Wu, Ji-Yuan Shi et al.· 0 citations
WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.
Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al.· 1 citation
KineBench is an IDM-free closed-loop benchmark for EWMs, built upon an explicit kinematic grounding pipeline, and reveals task-complexity-bounded nonlinear scaling in embodied video generation, providing empirical guidance for future data-scaling strategies.
Zeyu Liu, Zhang-Zhe Zhu, Yang Zhang et al.· arXiv.org· 2 citations
Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation, highlighting the importance of grounding action representations in executable control structure and environment-conditioned visual change.
Yuecheng Xu, Tong Yang, Jingkai Jia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.