Vision-Centric World Models for Embodied Robots: Representations, Predictive Interfaces, and Evaluation
Abstract
Embodied robots need more than a description of the current image: they must estimate how the scene may change under motion, contact, and partial observation. Vision-centric world models provide this predictive layer, but they expose it through different state interfaces. This paper organizes the literature into four families according to the state available to planning: future observations, compact latent states, geometry-structured maps, and persistent entities or relations. The comparison examines how each state is formed, advanced, and queried; where inference and planning costs arise; and which errors matter in physical use. Representative methods show that observation prediction retains interpretable appearance but makes repeated rollout expensive. Latent dynamics reduce that cost while risking the loss of contact-scale variables. Geometric states support pose, clearance, and occupancy queries, although their reliability depends on calibration and timely updates. Entity-relational states preserve object identity and task relations, yet they remain vulnerable to binding errors under occlusion. Evaluation is therefore linked to the exposed state rather than to a single generic score. The resulting framework clarifies where temporal prediction, persistent geometry, and semantic identity complement one another in embodied planning.