Skip to content
Preprint

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

Aug 2026 · 1 citation · 52 references
Computer Science

TL;DR

This work introduces SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model, and shows that structured future supervision and direct future-stream access yield higher trajectory planning scores.

Abstract

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.

View source

Similar papers

Preprint Sep 2026

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as inp...

Jin-Yang Wang, Shi-Wei Li, Jun-Jian Wang et al. · 0 citations
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 4 citations
Preprint Sep 2026

FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution,...

Jie Wu, Yu-Zhi Huang, Jun-Qi Liu et al. · 0 citations
Preprint Sep 2026

MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual struc...

Yi-Guang Yang, Jian-Kun Peng, Xiao-Ming Wang et al. · 0 citations
Preprint Aug 2026

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

Adaptive-WAM is introduced, a quality-aware multi-exit planner built on a Wan2.2-5B backbone that avoids the iterative classifier-free denoising loop and VAE decoding required for future-video synthesis, while dynamically allocating backbone depth according to trajectory quality.

Sining Ang, Yuguang Yang, Yan Wang · 2 citations
Preprint Sep 2026

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending...

Cong Wei, Xuan-Chi Ren, Bryan Chu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.