Application of Multimodal Fusion Based on Sensors and Machine Vision in Autonomous Driving
Abstract
Autonomous driving has become a transformative technology poised to reshape modern transportation systems. This paper explores multimodal fusion techniques that integrate various sensors with machine vision for autonomous driving. We examine the integration of various sensor modalities, including cameras, LiDAR, and millimeter-wave radar, alongside advanced machine vision algorithms such as YOLO, Faster R-CNN, Point Pillars, and MVX-Net for environment perception. This work addresses major challenges in sensor fusion, including data synchronization, coordinate transformation, real-time computation and conflict resolution of heterogeneous sensor data. We systematically analyze three typical fusion architectures: data-level, feature-level and decision-level fusion, and compare their performance in information retention, computational efficiency and system robustness. Through representative application cases in object detection and classification, high-precision localization and mapping, and decision-making and path planning, we demonstrate how multimodal fusion significantly enhances the robustness, accuracy, and reliability of autonomous vehicle perception systems. The paper further discusses current limitations including computational overhead, adverse-weather robustness, and lack of standardized evaluation, and outlines future directions such as end-to-end learning, 4D radar integration, and V2X-enabled cooperative perception. The results prove that reliable multimodal fusion is a core prerequisite for realizing safe and stable autonomous driving.