Skip to content
Open access

Self-Supervised 3D Point Cloud Detection with Depthwise Separable Masked Autoencoders

Sep 2026 · Remote Sensing · 0 citations · 12 references

Abstract

Three-dimensional object detection from LiDAR point clouds is essential for autonomous driving, yet existing methods typically rely on costly and extensively annotated datasets. Self-supervised masked autoencoders (MAE) provide a promising alternative, but effectively capturing local geometric structures while maintaining computational efficiency remains challenging. This paper presents DS-MAE, a depthwise separable masked autoencoder for self-supervised 3D point cloud detection. DS-MAE introduces a pyramidal transformer encoder with depthwise separable attention to enhance local geometric feature extraction while reducing computational cost. A depthwise separable generative decoder is further designed for multi-scale masked feature reconstruction, while a density-aware reconstruction loss accounts for the non-uniform point density of LiDAR observations. Experiments on the KITTI, Waymo Open Dataset (WOD), and a small-scale subset of the ONCE dataset demonstrate the effectiveness of DS-MAE, achieving 67.95%, 67.70%, and 58.92% mAP, respectively. On WOD, the reported performance is achieved using only 20% of the labeled training data, demonstrating the effectiveness of DS-MAE under limited supervision.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.