Skip to content

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Jul 2026 · arXiv.org · Vol abs/2607.13028 · 2 citations · 54 references
Computer Science

TL;DR

TerraZero is a procedural driving simulator and self-play training stack that meets the goals of reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.

Abstract

Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack that meets these goals. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per second on a single server-grade GPU, far faster than existing object-level simulators, while keeping fidelity lighter single-agent systems omit: heterogeneous agents, multiple dynamics models, and full traffic-rule enforcement. TerraZero uses logged data only as a source of real-world map geometry, populating each map with randomized rule-based road users and signal controllers and randomizing agent dynamics, rewards, and sizes per episode, so one map yields an effectively unbounded set of scenarios. Every reported policy trains from scratch by reinforcement learning alone, with zero human demonstrations, no imitation, no logged trajectories, and no fallback planner at inference, on a compute-efficient self-play recipe scaled across GPUs. The policies generalize zero-shot across cities and datasets, including emergent left-hand-traffic driving without explicit supervision. As an ego policy, a single checkpoint is, to our knowledge, the first fully learned policy to top both val14 and the interactive long-tail InterPlan suite. On Waymo Open Sim Agents realism the same recipe outperforms other demonstration-free methods and is competitive with the strongest reference-anchored self-play method. One stack serves both roles: state-of-the-art demonstration-free driving policies across dynamics for cars and trucks, and sim agents that jointly control vehicles, pedestrians, and cyclists.

View source

Similar papers

Preprint Aug 2026

Scaling Curriculum Learning For Autonomous Driving

CL4AD is presented, the first integration of curriculum learning into batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem, and utility functions that shape curricula based on success rates and the realism of the agent's behavior are introduced, in addition t...

Cevahir Koprulu, D. Paz, Feng Tao et al. · 1 citation
#machine learning Preprint Sep 2026

DashVMC: Real-Time Discrete World Model Control in Geometry Dash

World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Da...

Florent Tariolle, Florian Yger · 0 citations
Preprint Sep 2026

All You Need Is Low Fidelity: Zero-Shot Sim-to-Real of Learned Robotic Fish Control

Complex tasks for underwater robots remain limited by the capabilities of their controllers. Learning a better one for a soft, underactuated robotic fish trades simulator cost against fidelity. We show that an intentionally low-fidelity simulator is enough: a stateless, quasi-steady fluid model with no wake and no adde...

Liam Maloney, Simon Ramchandani, M. Y. Michelis et al. · 0 citations
Preprint Sep 2026

Calibrate Once, Fly Any Team: Residual-Grounded Low-Fidelity Training for Cooperative Drone Swarms

Training multi-agent drone-swarm policies directly in high-fidelity (HF) rigid-body physics is accurate but computationally expensive. This cost scales poorly with team size, as each additional agent multiplies contact-resolution complexity and sharply raises the in-simulation crash rate. To address this, we propose a...

Maxim Mednikov, Oren Gal · 0 citations

Learning the Right Abstraction: Neural Reduced Dynamics for Complex Robot Control

A neural reduced dynamics framework is developed that separates the state the model propagates from what can be supplied as an input or recovered analytically, trains policies entirely inside the frozen learned model, and validates them back in the high-fidelity simulator.

Harry Zhang, Dan Negrut · 0 citations
#artificial intelligence Preprint Sep 2026

Imitation Learning for Autonomous Driving in CARLA

This work studies how much closed-loop driving competence a compact multimodal policy can acquire from offline demonstrations in the CARLA simulator, and reports offline metrics and distinguish measured results from qualitative closed-loop observations.

Jordy Kieto · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.