SRCDet: sparse fusion of surround-view radar and camera for 3D object detection
Abstract
Multimodal fusion of cameras and millimeter-wave radars is critical for robust all-weather object detection and ensuring vehicle safety in real-world autonomous driving vehicles. However, existing radar-camera fusion methods that rely on a unified Bird’s Eye View (BEV) representation often suffer from information loss and limited cross-modal interaction. To address these limitations, a query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion. The framework processes features in parallel across both BEV and Perspective View spaces, where deformable attention is employed to achieve dynamic cross-view alignment. By integrating radar attribute features, a local–global dual-branch query generation mechanism is designed to produce high-quality 3D detection proposals. Furthermore, a graph neural network-based cross-fusion module is introduced to model complex inter-feature relationships through a heterogeneous interaction graph. Extensive experiments on the OmniHD-Scenes and NuScenes datasets demonstrate that SRCDet achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.