Skip to content
Preprint

Overcoming Statistical Bias in Action-Controllable World Models

Aug 2026 · 5 citations · 46 references
Computer Science

TL;DR

Co is introduced, a Counterfactual Consistency framework to enhance action controllability through two complementary constraints: Multi-step counterfactual consistency constrains reference, inverse-action, and zero-action rollouts, while action-spatial counterfactual consistency enforces consistent predictions under mirrored scenes and transformed actions.

Abstract

Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable from visual inertia and recurring motion patterns alone. This creates a shortcut: models can fit the data by exploiting statistical biases without making their visible dynamics meaningfully depend on the action. As a result, different actions may produce similar futures, while motion may persist even under zero action. The key question is how to reduce reliance on statistical shortcuts from dominating action-conditioned prediction. We argue that action control requires more than injecting action features; it requires enforcing consistency under counterfactual changes to actions and observations. Based on this insight, we introduce CoCo, a Counterfactual Consistency framework to enhance action controllability through two complementary constraints. Multi-step counterfactual consistency constrains reference, inverse-action, and zero-action rollouts, while action-spatial counterfactual consistency enforces consistent predictions under mirrored scenes and transformed actions. Together, they reduce reliance on statistical shortcuts from substituting for action-dependent dynamics. We further introduce Action Response Consistency (ARC) and Drift Energy (DE) to assess action controllability, together with Mini-SSMB for same-state, multi-action counterfactual evaluation. On Mini-SSMB, our full model achieved ARC_inv of 0.412 and ARC_ref of 0.483, while reducing DE by 17.07% relative to the baseline. On VP2 visual planning, it achieves the highest average success rate among SOTA models, at 73.1%. Experiments on BAIR and RoboNet further show that these gains preserve video prediction quality and transfer across model settings.

View source

Similar papers

Preprint Aug 2026

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

This work introduces the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions, and establishes the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting repr...

Junlin Chen, Rui-Jie Wang, Jian-Xin Li · 0 citations
#artificial intelligence Preprint Sep 2026

OneWorld: Learning Consistent Physics Across Actions in World Models

Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying inco...

Ke-Jun He, Yi-Chen Ding, Bin Yang · 0 citations
Preprint Aug 2026

How Can Driving World Models Do Counterfactual Prediction?

Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what would have happened under an alternative ego action. In this paper, we identify a fundamental mismatch between this goal and direct action-conditioned prediction. The d...

Jiaru Zhang, C. Cui, Yi Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which raises a question prediction quality alone cannot ans...

Todd Y. Zhou, Daniella Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Dynamic Manipulation with World-Action Models via Counterfactual Planning

DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data and constructs a counterfactual observation that places a predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery beh...

Sunwoo Park, Won-Sang Lee, Seonghyun Jin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.