Skip to content
Review

ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

Experiments on RoboTwin2.0 show that confidence-guided selection improves post-training efficiency, while dense frame and patch weighting further enhances prediction quality and embodied trajectory consistency compared with scalar reward, progress, and judge-based scoring baselines.

Abstract

Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but their errors under new task and scene distributions are often concentrated in localized spatiotemporal regions such as robot arms, manipulated objects, contact areas, and occluded objects. This paper presents ConfAL-WM, a confidence-guided active learning framework for post-training embodied world models. Built upon EVAC, we attach a lightweight confidence probe to UNet decoder features and predict dense confidence maps in the latent space. These maps are aggregated into task-, frame-, and patch-level scores, enabling both efficient data selection and localized training enhancement. Our pipeline first retrains the confidence probe and warms up EVAC with a small subset of target-domain data, then performs task-level prescreening to allocate sampling budgets, and finally applies selected-data retraining with optional frame or patch weighted data enhancement. Experiments on RoboTwin2.0 show that confidence-guided selection improves post-training efficiency, while dense frame and patch weighting further enhances prediction quality and embodied trajectory consistency compared with scalar reward, progress, and judge-based scoring baselines. A quick visual overview of this work is available at https://ConfAL-WM.github.io.

View source

Similar papers

Preprint Sep 2026

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-traini...

Chen-Hao Zhang, Han-Yu Zhao, Hang Cheng et al. · 0 citations
Preprint Aug 2026

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.

Fan Yang, Yu-Ting Su, Xiaobo Wang et al. · 9 citations
#artificial intelligence Preprint Aug 2026

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

IMPACT is introduced, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting, which consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.

Rong-Ze Tang, Jianjie Fang, Zhao-Lu Wang et al. · 1 citation
Preprint Aug 2026

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions, and is a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space.

Ruiteng Zhao, Zheng-Shen Zhang, Yue Su et al. · 2 citations
Preprint Sep 2026

Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Im...

Ke-Jia Hu, Wen Zhai, Bo-Tong Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.