Skip to content
Preprint

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

This work introduces the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions, and establishes the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation.

Abstract

Reinforcement-learning models commonly predict complete future states, observations, or feature occupancies, even though action selection depends only on differences between the consequences of candidate actions. As a result, these models may devote substantial statistical and representational capacity to high-dimensional phenomena that evolve independently of the agent's current choice. We introduce the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions. Its canonical centered representation removes this common component while preserving every pairwise action comparison expressible by the modeled reward family. The implemented model learns these action-dependent effects directly from synchronized counterfactual rollouts, so shared stochastic dynamics cancel before function approximation rather than after complete futures have been predicted. We establish the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation. Controlled experiments in physics-based environments provide initial evidence for these properties: direct effect learning suppresses action-independent variation, supports previously unseen reward queries, and improves action ranking relative to models trained to predict absolute futures.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which raises a question prediction quality alone cannot ans...

Todd Y. Zhou, Daniella Zhang · 0 citations
Preprint Aug 2026

Overcoming Statistical Bias in Action-Controllable World Models

Co is introduced, a Counterfactual Consistency framework to enhance action controllability through two complementary constraints: Multi-step counterfactual consistency constrains reference, inverse-action, and zero-action rollouts, while action-spatial counterfactual consistency enforces consistent predictions under mi...

Yu-Hong Shi, Zhenhao Chu, Jie Wei et al. · 5 citations
Preprint Aug 2026

How Can Driving World Models Do Counterfactual Prediction?

Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what would have happened under an alternative ego action. In this paper, we identify a fundamental mismatch between this goal and direct action-conditioned prediction. The d...

Jiaru Zhang, C. Cui, Yi Xu et al. · 0 citations
Preprint Aug 2026

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

The irreducible action-specific prediction error of future models that do not condition on the candidate action is characterized, conditions under which a world-action joint can recover an interventional forward model are identified, and an environment family is constructed in which every observational learner has posi...

Yu Yang · 1 citation
Preprint Aug 2026

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

This work proposes a compatibility prediction Latent World Model for robot navigation that predicts action-conditioned latent feature compatibility rather than reconstructing observations and demonstrates how the learned world model can supervise policy learning from unlabeled video data and improve policies through re...

Zengmao Wang, Wei Gao, Shuhan Shen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.