Skip to content

DriftWorld: Fast World Modeling through Drifting

Jul 2026 · arXiv.org · Vol abs/2607.15065 · 0 citations · 52 references
Computer Science

Abstract

Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. This creates a bottleneck for diffusion-based world models: multistep sampling makes each rollout expensive, limiting large-scale action search at inference time. We introduce DriftWorld, an action-conditioned world model based on drifting generative models. Rather than denoising iteratively at inference, DriftWorld learns an action-conditioned drift during training, allowing it to generate future frames from the current observation and a candidate action sequence in a single forward pass at 30+ fps, which is 17x faster on average than diffusion based baselines. We evaluate DriftWorld on standard vision-based robotic manipulation benchmarks, including Bridge-V2, RT-1, Language Table, Push-T, and Robomimic. By producing rollouts that are both accurate and fast, DriftWorld achieves state-of-the-art decision-making performance with far less inference time than diffusion-based world model baselines. Beyond online control, DriftWorld can also serve as an offline simulator for ranking real-world robot policies, with rollout-based scores correlating with ground truth at up to 0.99. These results show that drifting models are a strong fit for robot world modeling, where fast, high-quality imagination directly supports planning and policy evaluation.

View source

Similar papers

Jul 2026

Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

Temporal Policy is introduced, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem and bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control.

Dylan Miller, Martin Jägersand · 0 citations
Preprint Aug 2026

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.

Chunkai Yang, An-Dong Yang, Di Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LePlanner: An Iterative Amortized Controller For World Models

World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet planning in these latent spaces typically relies on one of two costly approaches. Search-based planners such as CEM, MPPI, and iCEM optimize action sequences through many pr...

Saksham Bansal, Om Naphade, Chayan Aggarwal et al. · 0 citations
Preprint Aug 2026

CoDrift: Compositional Drifting for Offline Reinforcement Learning

This work proposes CoDrift, a compositional framework for one-step generative policy learning that combines three objective-level fields into a unified policy field that compares favorably with state-of-the-art methods and achieves the best average rank in both settings.

Xiewei Ni, Ruo-Feng Mei, Xiang-Yu Xu · 0 citations
Preprint Aug 2026

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

WorldCycle is introduced, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency rewar...

Bohai Gu, Yueyang Yuan, Tai-Yi Wu et al. · 2 citations
Conference Open access Sep 2026

Context-Adaptive World Models for Visual Out-of-Distribution Generalization

This thesis is developing a context-adaptive Mamba-based world model that infers K context vectors from short pixel-trajectory windows, and shows that its suboptimality is bounded by √K times the context-inference error along a Lipschitz chain through FiLM, the SSM, and the decoder.

Shubham Subhnil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.