World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution,...
Jie Wu, Yu-Zhi Huang, Jun-Qi Liu et al.· 0 citations
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term t...
Yuzhi Huang, Wei Bu, Ziyi Xiong et al.· 1 citation
TAU-Bench is introduced, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding, and shows that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic re...
Kepeng Yang, Dong-Xuan Liu, Rongxin Gao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.