Skip to content

TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation

Jul 2026 · arXiv.org · Vol abs/2607.10762 · 1 citation · 36 references
Computer Science

TL;DR

TOLiD is proposed, a self-supervised pretraining method for LiDAR representation learning that addresses the gap between dense ViT token representations and sparse 3D encoders by coupling a LiDAR backbone with a student Vision Transformer initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations.

Abstract

Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding. However, existing distillation pipelines typically treat the VFM as a frozen feature source and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. We propose TOLiD, a self-supervised pretraining method for LiDAR representation learning that addresses this gap by coupling a LiDAR backbone with a student Vision Transformer (ViT) initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations. TOLiD converts the set of point features within each image patch frustum into a token using Frustum Pooling followed by Frustum Attention, and performs token-level distillation with visibility masking. For LiDAR-only deployment, we lift token features back to per-point representations using masked bilinear sampling to avoid patches that have limited LiDAR points. We extensively evaluate TOLiD on five heterogeneous LiDAR datasets and four cross-sensor adaptation pairs, demonstrating improved transfer with frozen backbones and lightweight heads.

View source

Similar papers

Preprint Sep 2026

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...

Yu-Hang Han, Youngseok Jang, Seungwon Roh et al. · 0 citations
Preprint Aug 2026

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

CVSD-Reg is proposed, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations and generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference.

Eunsoo Im, Junghun Suh, Gyeonggwan Lee et al. · 0 citations
Preprint Aug 2026

Vernata: Self-Supervised Learning of LiDAR Point Representations

Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...

Oliver Lemke, Alexander Liniger, Abel Gawel et al. · 0 citations
Preprint Aug 2026

Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models

The Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2, achieves state-of-the-art performance and ablation studies confirm the importance of teacher-student semantic compatibility.

Qifeng Zhang, Ting Xiang, Ze-Yu Bai et al. · 0 citations
Jul 2026

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

This work proposes Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement and introduces Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discret...

Rui Tang, Wentao Yang, Peirong Zhang et al. · 0 citations
Preprint Sep 2026

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independ...

Yang Zhou, Jiuhong Xiao, Shi-Zhao Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.