Preprint
Aug 2026
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
This work introduces SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model, and shows that structured future supervision and direct future-stream access yield higher trajectory planning scores.
Yibo Yuan, Jiacheng Fu, Jiangtong Zhu et al.
· 1 citation