Two-Stage 3D Object Detection Architecture via Integrated Attention Mechanism and Voxel Aggregation
Abstract
In the realm of autonomous driving, 3D object detection based on LiDAR point clouds has emerged as a pivotal technology. To enhance the accuracy and efficiency of 3D object detection, this paper introduces Spatial-Channel Attention Guided with Gumbel Subset Sampling and Context Fusion RCNN (SCAGCF-RCNN), a two-stage detection framework that integrates spatial-channel feature enhancement with hierarchical voxel aggregation. In the first stage, SCAGCF-RCNN leverages a sparse convolutional network to extract multi-scale 3D point cloud features, which are then projected onto a Bird's Eye View (BEV) representation. The Spatial and Channel Synergistic Attention (SCSA) module is employed to capture contextual information and enhance feature representation by synergistically combining spatial and channel self-attention mechanisms. In the second stage, the Gumbel Subset Sampling with Context Fusion (GSSCF) is utilized to downsample the original point cloud, selecting key points that retain critical geometric structure information. The context fusion module further optimizes these key points by converging multi-level features. Finally, the Enhanced Extended Voxel Set Abstraction (EEVSA) module fuses multi-scale voxel features with 2D BEV features and key point features, generating refined proposal features for accurate detection. Extensive experiments on the KITTI and Waymo datasets demonstrate that SCAGCF-RCNN achieves competitive accuracy and robustness, particularly in challenging detection scenarios.