We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances...
Min-Xing Li, Ming-Hao Han, Wei-Zhi Zhao et al.· 0 citations
WorldExam is introduced, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity, which supports unified evaluation of camera-, action-, and language-driven model paradigms.
Yuxue Yang, Shu-Yao Shang, Jiahe Wang et al.· 2 citations
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Mot...