A Survey of Deep Learning-Driven Multi-Modal Perception Fusion for Autonomous Driving
This paper presents a comprehensive survey on deep learning based multi-modal perception fusion frameworks for autonomous driving systems. The traditional single-modal solutions have intrinsic limitations in robustness against harsh weather and dynamic environments, and the integration of complementary sensors, such as cameras, LiDAR, and millimeter-wave radar, provides a critical pathway to high-level automation. We provide a systematic review of the basic fusion paradigms, classifying them into early, late and deep (feature-level) architectures, and evaluating their trade-offs in terms of computational overhead, architectural modularity and joint feature learning. Special focus is given to state-of-the-art architectures, comparing the high representational power of the Tensor Fusion Networks (TFN) to the computationally-efficient Multi-modal Circulant Fusion (MCF) and context-aware Transformer frameworks. This study places real-world deployment challenges in the context of algorithmic paradigms, investigating the friction between “black-box” deep learning models and stringent functions. Besides algorithmic paradigms, this work puts real-world deployment challenges into context and discusses the friction between ``black-box'' deep learning models and rigorous functional safety standards (e.g., ISO 26262), and energy constraints of new energy vehicles. Finally, we explore the paradigm shift to "vehicle-road synergy" (V2X) infrastructure as a critical mechanism for providing the safety redundancy and edge-computing capabilities needed for fully reliable, next-generation autonomous driving.