We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances...
Min-Xing Li, Ming-Hao Han, Wei-Zhi Zhao et al.· 0 citations
WorldExam is introduced, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity, which supports unified evaluation of camera-, action-, and language-driven model paradigms.
Yuxue Yang, Shu-Yao Shang, Jiahe Wang et al.· 2 citations
CHOREO, a framework for training-free composition of heterogeneous humanoid skills, demonstrates that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.
Zi-Yi Sun, Jing-Wen Chen, Yu-Xin Wang et al.· 1 citation
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Mot...