Class-Enhanced Multi-Sampling and Multi-Level Graph Attention for 3-D Object Detection
Abstract
Three-dimensional object detection is a key component of the perception module in autonomous driving systems. Compared to camera images, LiDAR point clouds provide richer spatial information, such as detailed structural and geometric cues of objects. However, existing 3D object detection methods face two major challenges: 1) loss of critical information due to downsampling or convolution operations, leading to missed detections, and 2) insufficient exploitation of contextual information around objects, resulting in inaccurate bounding box regression. These problems are particularly severe when detecting sparse point cloud objects that are distant or occluded. To address these issues, we propose two novel modules applicable to both point-based and voxel-based detection frameworks: the Class-Enhanced Multi-Sampling (CEMS) module and the Multi-Level Graph Attention (MLGA) module. The CEMS module employs an iterative class-aware sampling and semantic interpolation strategy to achieve more precise downsampling of critical information, while the MLGA module integrates both intra-layer and cross-layer graph attention to capture local geometric structures around objects as well as inter-object relationships, thereby enhancing the representation of key features. Comprehensive evaluations on multiple datasets, including ONCE, Waymo, and nuScenes, demonstrate that the proposed modules can consistently improve the detection performance across various network architectures while maintaining high inference efficiency.