ARSTAG is presented, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data and shows that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size.
Abstract
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.
Meng-Zhao Jia, Yang Lin, Xi-Xin Zhang et al.· 10 citations· ⚡1
Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting,...
Hao-Ran Lang, Hao-Tao Lu, Shi-Yu Sang et al.· 1 citation
This work presents RobotMover, a complete learning-based system for large-object manipulation that leverages human–object interaction demonstrations to train robot control policies and achieves strong performance in terms of capability, robustness, and controllability, outperforming both learned and teleoperation basel...
Tian-Yu Li, Joanne Truong, Tsung-Yen Yang et al.· IEEE Transactions on robotic...· 2 citations
DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.
Makoto Sato, T. Matsushima, Yutaka Matsuo et al.· 1 citation
Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop co...
He Zhu, Lu-Sen Zhao, Kwan-Man Cheng et al.· 0 citations
This work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration, using an object-centric relational program representation.
Yu-Yao Liu, Jia-Yuan Mao, David Hsu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.