Radar--Camera Fusion Vehicle Detection Algorithm With Adaptive Constraints on Point Cloud Depth and RCS
Abstract
In traffic scenes, the degradation of radar point cloud angle-of-arrival estimation accuracy with increasing distance leads to reduced discriminability and spatial response dispersion of radar–camera fusion features in bird’s-eye view (BEV) space. To address this issue, an improved Astra-HGSFusion radar–camera fusion vehicle detection algorithm is proposed. The method integrates an adaptive radar hybrid generation module (ARHGM) that jointly constrains virtual point cloud generation by target depth and radar cross-section (RCS), a multipooling feature encoding mechanism, and a shareable multisemantic spatial attention mechanism. Specifically, a hybrid sampling strategy combining depth-adaptive Gaussian kernels and RCS-adaptive sampling is introduced, with semantic information utilized for auxiliary generation to enhance the virtual point cloud. The multipooling feature encoding effectively preserves fine-grained spatial information within Pillars, while a multisemantic spatial attention module is incorporated into the radar branch feature pyramid network to strengthen target feature responses in the BEV representation. Consequently, the proposed algorithm demonstrates superior detection performance in sparse radar scenarios. Experimental validation on the View-of-Delft (VOD) and TJ4DRadSet datasets shows that the proposed method achieves mean average precision (mAP) scores of 81.07% and 60.96% in the driving region and the fully annotated region, respectively, representing improvements of 2.20% and 3.19% over the HGSFusion baseline. Furthermore, BEV mAP and 3-D mAP are enhanced by 3.40% and 3.70%, respectively. Simulation results demonstrate that the algorithm effectively mitigates missed detections and localization errors for distant and small-scale targets, providing valuable theoretical reference for multimodal point cloud-image perception in intelligent transportation systems.