Skip to content

Discriminative World Models for Web Agents

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

Predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions, is introduced and outperforms world models trained with supervised next-state prediction.

Abstract

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates to accurately score them. To address this, we introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions. We train these models using a branching web-agent dataset derived from WebArena Go-Browse trajectories, where every decision point contains multiple alternative actions and their resulting states. Experiments on our held-out predicted-state matching benchmark show that our approach outperforms world models trained with supervised next-state prediction. We further show that our approach improves PRM-style action ranking on WebPRMBench compared with action-only PRMs and PRMs augmented with supervised-next-state world models. Finally, on WebArena-Lite, using our world model for test-time action selection improves end-to-end task success. Our project page is available at: https://dhruvpendharkar.github.io/dwm/.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training

Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observa...

Xin-Yu Che, Hang Yan, Yan-Chen Liu et al. · 0 citations
Preprint Sep 2026

Recommendation World Models for Future-State Control

Sequential recommendation optimizes which items to rank, while each displayed slate also shapes subsequent feedback and user state. We study how a trained ranker can support decisions about these future consequences. We introduce UA-TWM, a utility-anchored world-model interface that constructs nearby slate actions, est...

Jin-Feng Xu, Zhe-Yu Chen, Zi-Yue Peng et al. · 0 citations
Preprint Aug 2026

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

This work proposes a compatibility prediction Latent World Model for robot navigation that predicts action-conditioned latent feature compatibility rather than reconstructing observations and demonstrates how the learned world model can supervise policy learning from unlabeled video data and improve policies through re...

Zengmao Wang, Wei Gao, Shuhan Shen · 0 citations
Preprint Sep 2026

The Planning Limits of Latent World Models

World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this ques...

A. Alrasheed, Basim Azam, Naveed Akhtar · 0 citations
Preprint Sep 2026

From World Models to World Action Models: Rethinking Next-State Prediction

Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets,...

Ting-Yu Yuan, Zi-Ming Ji, Biao-Liang Guan et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.