TOLiD is proposed, a self-supervised pretraining method for LiDAR representation learning that addresses the gap between dense ViT token representations and sparse 3D encoders by coupling a LiDAR backbone with a student Vision Transformer initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations.
Abstract
Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding. However, existing distillation pipelines typically treat the VFM as a frozen feature source and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. We propose TOLiD, a self-supervised pretraining method for LiDAR representation learning that addresses this gap by coupling a LiDAR backbone with a student Vision Transformer (ViT) initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations. TOLiD converts the set of point features within each image patch frustum into a token using Frustum Pooling followed by Frustum Attention, and performs token-level distillation with visibility masking. For LiDAR-only deployment, we lift token features back to per-point representations using masked bilinear sampling to avoid patches that have limited LiDAR points. We extensively evaluate TOLiD on five heterogeneous LiDAR datasets and four cross-sensor adaptation pairs, demonstrating improved transfer with frozen backbones and lightweight heads.
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...
Yu-Hang Han, Youngseok Jang, Seungwon Roh et al.· 0 citations
CVSD-Reg is proposed, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations and generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference.
Eunsoo Im, Junghun Suh, Gyeonggwan Lee et al.· 0 citations
Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...
Oliver Lemke, Alexander Liniger, Abel Gawel et al.· 0 citations
The Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2, achieves state-of-the-art performance and ablation studies confirm the importance of teacher-student semantic compatibility.
Qifeng Zhang, Ting Xiang, Ze-Yu Bai et al.· 0 citations
This work proposes Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement and introduces Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discret...
Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independ...
Yang Zhou, Jiuhong Xiao, Shi-Zhao Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.