Aug 2026· 5 citations· ⚡ 1 influential· 53 references
Computer Science
TL;DR
FACT is introduced, a causal World-Action Model that predicts future video and task progress conditioned on the executed action, and allows failure rollouts to supervise action consequences, turning bad actions into valid future targets rather than being discarded.
Abstract
Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on the future-prediction ability of video models, many WAMs generate future videos and recover actions with inverse-dynamics models, or use these predicted videos as goal conditions for action generation. In both cases, the world model is trained mostly on successful demonstrations and has little reason to predict the consequences of bad actions. We introduce FACT, a causal World-Action Model that predicts future video and task progress conditioned on the executed action. This action-conditioned interface allows failure rollouts to supervise action consequences, turning bad actions into valid future targets rather than being discarded. Failure-aware training makes the progress predictor aware of both successful and failed action outcomes, which can optionally be used to score sampled action candidates at inference. Extensive experiments on simulation and real-world bimanual manipulation tasks show that FACT outperforms many existing baselines, improves as failure data are incorporated into training, and reduces success-biased future hallucination under bad actions. See more details at https://fact-wam.github.io/
These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts that achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.
Qiwen Gu, Jifan Li, Bing-Jie Gao et al.· 0 citations
ForeWAM is proposed, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos, and demonstrates that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating futu...
Jiakai Huang, Zhongbo Wu, Siyu Xu et al.· 3 citations
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives,...
SelfWAM is introduced, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion.
Bikang Pan, Fan Liu, Haotao Lu et al.· 5 citations
This work proposes a compatibility prediction Latent World Model for robot navigation that predicts action-conditioned latent feature compatibility rather than reconstructing observations and demonstrates how the learned world model can supervise policy learning from unlabeled video data and improve policies through re...
WorldEcho is introduced, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment, and WorldSync is proposed, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alig...