Skip to content
Preprint

The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts that achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.

Abstract

Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear action decoding easier. Observed future changes provide additional action information beyond the present, and linearly readable action information is spatially concentrated. These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts. ARPS uses a horizon-conditioned state predictor to aggregate intermediate video features into a compact state that supplies all visual context to the action expert. Future-representation supervision trains different parts of this state to predict visual representations at different future times, together with their changes relative to the present. At inference, the supervision branch is removed, and the action expert only uses the learned predictive state computed from current observations. Controlled ablations show that future supervision substantially improves generalization under distribution shift. ARPS achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.

View source

Similar papers

Preprint Aug 2026

FACT: Failure-Aware Causal Training for World-Action Models

FACT is introduced, a causal World-Action Model that predicts future video and task progress conditioned on the executed action, and allows failure rollouts to supervise action consequences, turning bad actions into valid future targets rather than being discarded.

Quanquan Peng, Yutong Liang, Rui Yan et al. · 5 citations · ⚡1
Preprint Aug 2026

Foresight Without Seeing: Latent Futures for World Action Models

ForeWAM is proposed, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos, and demonstrates that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating futu...

Jiakai Huang, Zhongbo Wu, Siyu Xu et al. · 3 citations
Preprint Sep 2026

From World Models to World Action Models: Rethinking Next-State Prediction

CF-WAM is proposed, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections.

Ting-Yu Yuan, Zi-Ming Ji, Biao-Liang Guan et al. · 1 citation
#machine learning Preprint Sep 2026

MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies

Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared g...

Jing-Qi Wang, Yan Wang · 0 citations
Preprint Aug 2026

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives,...

Xiang Li, Yu-Peng Zheng, Song-En Gu et al. · 1 citation · ⚡1
Preprint Oct 2026

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT reco...

Chensheng Peng, Wen-Hao Ding, Ran Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.