Skip to content

PhiZero: A World Model Built Around Physical Language

Jul 2026 · arXiv.org · Vol abs/2607.28624 · 1 citation · 105 references
Computer Science

Abstract

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans'ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.

View source

Similar papers

Jul 2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

Jiaxin Bai, Jia–Jie Xiong · 0 citations
Preprint Sep 2026

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Qualitative examples show the PhysBrain 1.5 model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

DeepCybo Team, Yue Bin, Hai-Peng Cao et al. · 0 citations
Preprint Aug 2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

World Tokens is an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation and is highly competitive on LIBERO, attains the best reported averages on SIMPLER, and substantially improves real-world R1 Pro success over a matched...

Qu Tang, Benhui Zhuang, Bo Yuan et al. · 1 citation · ⚡1
Preprint Aug 2026

Code World Model: Coding Agent as World Brain

MiniMax-H3 follows proxy-based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics, demonstrating the potential of combining code for persistent world evolution with video models for flexible visual realization.

Yiwen Chen, Guo-Sheng Lin, Chi Zhang · 7 citations
Preprint Sep 2026

Programmable World Model

Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visu...

Zheng-Hui Huang, Gui-Xu Lin, Jia-Cheng Lin et al. · 5 citations · ⚡1
Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.

Xuan Yao, Junyu Gao, Chang-Sheng Xu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.