Experiments on RoboTwin2.0 show that confidence-guided selection improves post-training efficiency, while dense frame and patch weighting further enhances prediction quality and embodied trajectory consistency compared with scalar reward, progress, and judge-based scoring baselines.
Abstract
Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but their errors under new task and scene distributions are often concentrated in localized spatiotemporal regions such as robot arms, manipulated objects, contact areas, and occluded objects. This paper presents ConfAL-WM, a confidence-guided active learning framework for post-training embodied world models. Built upon EVAC, we attach a lightweight confidence probe to UNet decoder features and predict dense confidence maps in the latent space. These maps are aggregated into task-, frame-, and patch-level scores, enabling both efficient data selection and localized training enhancement. Our pipeline first retrains the confidence probe and warms up EVAC with a small subset of target-domain data, then performs task-level prescreening to allocate sampling budgets, and finally applies selected-data retraining with optional frame or patch weighted data enhancement. Experiments on RoboTwin2.0 show that confidence-guided selection improves post-training efficiency, while dense frame and patch weighting further enhances prediction quality and embodied trajectory consistency compared with scalar reward, progress, and judge-based scoring baselines. A quick visual overview of this work is available at https://ConfAL-WM.github.io.
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-traini...
Chen-Hao Zhang, Han-Yu Zhao, Hang Cheng et al.· 0 citations
LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.
Fan Yang, Yu-Ting Su, Xiaobo Wang et al.· 9 citations
IMPACT is introduced, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting, which consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
Rong-Ze Tang, Jianjie Fang, Zhao-Lu Wang et al.· 1 citation
AcrossWAM1.0 contributes an auditable software and evaluation boundary for compact latent world-action policies, distinct from LaWAM's original latent-subgoal contribution.
SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions, and is a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space.
Ruiteng Zhao, Zheng-Shen Zhang, Yue Su et al.· 2 citations
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Im...
Ke-Jia Hu, Wen Zhai, Bo-Tong Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.