A review of how modular end-to-end architectures evolve with respect to scene representation and planning design and highlights recurring challenges, including error propagation across tasks, robustness under distribution shift, long-horizon reasoning, and evaluation under closed-loop interaction.
Abstract
Modular end-to-end autonomous driving has emerged as a middle ground between classical modular stacks and opaque end-to-end pipelines, aiming to retain the trainability of end-to-end learning while improving interpretability through structured intermediate representations. This study presents a review of how modular end-to-end architectures evolve with respect to 1) scene representation and 2) planning design. We organize representative systems by their dominant scene representation — dense bird’s-eye-view (BEV) grids, vectorized/map-centric representations, and sparse token-based representations — and discuss how each choice shapes the design of prediction and planning heads, computational cost, and safety-critical failure modes. In parallel, we survey planning strategies used in these architectures, contrasting deterministic single-plan selection with probabilistic planning that explicitly models multi-modal futures and risk. Through a comparison of recent benchmarks and survey literature, we highlight recurring challenges, including error propagation across tasks, robustness under distribution shift, long-horizon reasoning, and evaluation under closed-loop interaction. Finally, we summarize open research directions for modular end-to-end driving, including scalable sparse representations, uncertainty-aware planning, representation–planning co-design, verification, and the integration of foundation and world models.
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by...
Yue-Ting Zhu, Shao-Yu Chen, Yue-Hao Song et al.· 0 citations
End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imitation learning, privil...
Yanchen Guan, Xing-Chen Liu, Bin Rao et al.· 0 citations
We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cro...
Zhao-Wei Lu, Li-Guo Zhou, Yu-Jie Guo et al.· 0 citations
PAVER, Planning-Aligned BEV Encoder Pretraining is introduced, where from a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions, preserving the downstream architecture and camera-only inference.
This work proposes FeasibleFlow, a one-step end-to-end generative framework that jointly transports a configuration-space feasibility field and multimodal ego trajectories and introduces the Anchor-relative ranker (ARR) and Pareto-ReinFlow to balance safety and progress in candidate selection and generation.
Xiang Li, Bi-Kun Wang, Qing Xu et al.· 0 citations
Autonomous driving requires more than recognizing what is present in a scene: a planner must determine how road structure, surrounding agents, and their motion states should influence a future maneuver. Existing learning-based planners can capture these influences through latent scene features and trajectory decoders,...
Zhi-Yuan Liu, Yuan-Xin Tian, Ze-Hong Ke et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.