Skip to content
Preprint

Code World Model: Coding Agent as World Brain

Aug 2026 · 7 citations · 91 references
Computer Science

TL;DR

MiniMax-H3 follows proxy-based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics, demonstrating the potential of combining code for persistent world evolution with video models for flexible visual realization.

Abstract

World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world evolution. This makes it difficult to maintain persistent consequences and support coherent, open-ended evolution. We introduce Code World Model, a framework that separates world evolution from visual realization by combining the reasoning and coding capabilities of language models with the generative priors of video models. A coding agent serves as the world brain, reasoning about events and their consequences and generating executable code to maintain persistent world state and perform rule-consistent evolution. To connect executable state with visual generation, we introduce a proxy representation that encodes frame-wise spatiotemporal constraints and is compiled into a proxy video, which conditions a video model to render high-fidelity visual observations. We further develop data pipelines for constructing aligned proxy-observation pairs from gameplay and real-world videos. After fine-tuning on paired gameplay data, MiniMax-H3 follows proxy-based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics. These results demonstrate the potential of combining code for persistent world evolution with video models for flexible visual realization, providing a new path toward open-ended world models.

View source

Similar papers

Jul 2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

Jiaxin Bai, Jia–Jie Xiong · 0 citations
Preprint Sep 2026

Programmable World Model

Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visu...

Zheng-Hui Huang, Gui-Xu Lin, Jia-Cheng Lin et al. · 5 citations · ⚡1
Preprint Aug 2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

World Tokens is an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation and is highly competitive on LIBERO, attains the best reported averages on SIMPLER, and substantially improves real-world R1 Pro success over a matched...

Qu Tang, Benhui Zhuang, Bo Yuan et al. · 1 citation · ⚡1
Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.

Xuan Yao, Junyu Gao, Chang-Sheng Xu · 0 citations
Review Sep 2026

WorldReward: Reward Modeling for Camera-Conditioned World Models

This work presents WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models, and introduces WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency,...

Yi-Bin Wang, Ze-Han Wang, Junshu Tang et al. · 0 citations
Jul 2026

PhiZero: A World Model Built Around Physical Language

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Mot...

Shuyao Shang, Yuqi Wang, Ruopeng Gao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.