Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts object dynamics using larger time steps than the source numerical simulator, which requires small integration steps to resolve rapid motion and prevent interpenetration. We evaluate WorldContact across 16 shopping-bag manipulation tasks. State-rollout measurements on a single H100 GPU show a $10\times$ speedup over the source simulator, excluding rendering and disk I/O. We use the generated data to fine-tune an existing vision-language-action policy and deploy it directly on a real robot. In bag lifting, the same policy achieves 65% single-attempt success when fine-tuned on source simulation data alone, compared with 95% when fine-tuned on the dataset expanded with WorldContact. These results support efficient data generation with WorldContact for robot policy adaptation.
Caoliwen Wang, Meng-Di Wang, Heng Zhang et al.· 0 citations
TailorCoPilot demonstrates a viable pathway to capture practice-based expertise, operationalizing it to support both generative AI advancements and human apprenticeship.
Yue Sun, Zhao-Hui Wang, Ruiyang Liu et al.· 0 citations
TAMF-VTON is presented, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions and outperforms state-of-the-art methods in both quantitative metrics and perceptual quality.
Jie Wang, Qian He, Gaofeng He et al.· arXiv.org· 0 citations
Agentic Real2Sim is introduced, a framework for generalized physical world modeling with vision-language agents, converting a real-world recording of object-robot interaction into a simulatable episodic twin which preserves observations, geometries, robot interactions, and object states.