OccM3D-SS: Bridging Sparse Supervision With Dense 3D Perception via Semi-Supervised Occupancy Learning
Abstract
Monocular 3D object detection faces inherent challenges in reconstructing accurate 3D spatial structures from 2D images due to the lack of depth cues. Existing paradigms either struggle with limited 3D representation capability, e.g. depth estimation pipeline, or rely heavily on fine-grained supervision annotations, e.g. occupancy-guided pipeline. Moreover, occupancy labels derived from sparse LiDAR signals are often unreliable, introducing noise into geometric supervision. In this work, we propose OccM3D-SS, a semi-supervised occupancy-guided framework that integrates occupancy learning with unlabeled sequential data to enhance geometric reasoning. Our framework employs a dual-branch architecture: the supervised branch enforces semantic and geometric perception using annotated data, while the unsupervised branch leverages occupancy constraint to enhance geometric representation capability. Furthermore, to address the label noise issue, we introduce a temporal point-cloud densification strategy by tracking and aggregating foreground points from multi-frames, thereby improving the quality of occupancy annotations significantly. The experimental results on the KITTI-3D and Waymo Open datasets demonstrate that our framework outperforms the baseline by a large margin, achieving state-of-the-art performance while reducing annotation dependency.