World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution,...
Jie Wu, Yu-Zhi Huang, Jun-Qi Liu et al.· 0 citations
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term t...
Yuzhi Huang, Wei Bu, Ziyi Xiong et al.· 1 citation
TAU-Bench is introduced, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding, and shows that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic re...
Kepeng Yang, Dong-Xuan Liu, Rongxin Gao et al.· 1 citation
It is shown that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode the authors term prediction-preserving sharpening, and proposed Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a...
Jing-Yan Jiang, Yaru Sun, Xiao Chen et al.· 0 citations
GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in co...