Skip to content
Preprint

GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

Aug 2026 · 1 citation · 44 references
Computer Science

TL;DR

GhostPoint is proposed, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation, and introduces a predictor-level supervision scheme on sampled voxels from generated neighborhoods.

Abstract

3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer to 3D detection remains weaker. We analyze recent SSL methods and find that most objectives are defined only on measured LiDAR returns from visible surfaces, leaving occluded and unobserved regions unconstrained. This visible-surface bias can be sufficient for point-wise prediction, but 3D detection requires robustness to missing structure. To address this gap, we propose GhostPoint, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation. In GhostPoint, an encoder processes observed returns, and an additional predictor infers neighborhood representations from observed context. In addition to standard encoder-level supervision, we introduce a predictor-level supervision scheme on sampled voxels from generated neighborhoods. Specifically, observed (visible/masked) voxels match teacher-encoder targets, while unobserved voxels match teacher-predictor hallucinations. This design encourages the learned representation to explicitly model structure beyond observed returns. Extensive evaluations on nuScenes and Waymo demonstrate that our method achieves state-of-the-art performance, consistently improving downstream 3D detection, especially under sparse scans and limited labels.

View source

Similar papers

Preprint Aug 2026

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither...

Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al. · 0 citations
Preprint Aug 2026

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

CVSD-Reg is proposed, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations and generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference.

Eunsoo Im, Junghun Suh, Gyeonggwan Lee et al. · 0 citations
Preprint Aug 2026

Vernata: Self-Supervised Learning of LiDAR Point Representations

Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...

Oliver Lemke, Alexander Liniger, Abel Gawel et al. · 0 citations
Preprint Sep 2026

Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge

Transferring the rich priors of large 2D foundation models to sparse 3D LiDAR remains challenging, as training native 3D foundation models at comparable scale is limited by data and annotation scarcity. We introduce a LiDAR-conditioned diffusion model trained on pseudo-labels from off-the-shelf 2D foundation models. Th...

Samed Doğan, Nico Leuze, Alfred Schöttl · 0 citations
Preprint Sep 2026

Unsupervised Point Cloud Registration via Training-Time Semantic Guidance

Unsupervised registration of large-scale LiDAR point clouds remains challenging due to the geometric ambiguity inherent in outdoor scenes, which degrades pseudo-label quality and leads to suboptimal convergence, particularly for sparse, low-resolution scans such as those from nuScenes. We reveal that registration model...

Kezheng Xiong, Shi-Yun Xu, Sheng Ao et al. · 0 citations
Open access Aug 2026

STG: Structured Topology of Gridpoints for Occluded Pedestrian Detection

Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose...

Tian Qiu, Jifeng Shen, Xin Zuo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.