Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and comple...
Yu-Zhou Wu, Long-Teng Fan, Zi-Meng Li et al.· 0 citations
The results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.
Mu-Chun Niu, Shuang Chen, Yu-Zhou Wu et al.· 0 citations
A system level acceleration strategy that reduces computation in both perception and action generation and compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision is proposed.
Method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data, is presented, demonstrating that the synthesized data substantially improve downstream WAM generalization.
Zexuan Yan, Yuzhou Wu, Yue Ma et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.