Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can e...
Zhan-Peng He, Joaquin Palacios, Zhang-Yu Wang et al.· 0 citations
Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformab...
You-Hui Wang, Yun-Zhu Li, Fei-Fei Li et al.· 0 citations
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use h...
Yi-Ze Liu, Huang Huang, Yi-Ning Hong et al.· 0 citations
A robot may lose sight of an object it must later retrieve, need to recall what a person demonstrated earlier, or track which steps of a task it has already completed. Current vision-language-action (VLA) policies often fail once the information needed for action disappears from the current observation, making memory c...
Yi-Ze Liu, Ke Wang, Mac Schwager et al.· 0 citations
A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.
Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu· Communications of the ACM· 14 citations· ⚡1
Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LOTA...
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metri...
Yun-Fei Ge, An-Bang Liu, Qineng Wang et al.· 0 citations
DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points to improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved.
TrAct is proposed, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction, enabling more accurate world modeling and stronger robot generalization.
Zhi-Hang Cao, Howard Ji, Kevin Zhang et al.· 3 citations
FL-MAESTRO is proposed, a multi-agent orchestrator that makes the joint runtime FL decision directly through three specialist LLM agents, one per decision dimension, and matches the accuracy of the strongest energy-aware baseline while cutting wasted round energy from over a third to near zero.
Jia-Jun Wu, Zi-Rui Wang, Jia-Yu Zhou et al.· 0 citations
WebWorld is presented, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation.
Jia-Jun Wu, Jian Yang, Ya-Xin Du et al.· 0 citations
Masked Visual Actions is introduced, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video that supports inverse modeling by synthesizing robot motion from desired object motion.
Hadi Alzayer, Wenlong Huang, Haonan Chen et al.· arXiv.org· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.