Skip to content
Preprint

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

CVSD-Reg is proposed, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations and generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference.

Abstract

Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3 student learns from a frozen DINOv2 teacher through contrastive distillation and spherical-manifold alignment, which preserves the hyperspherical geometry of the teacher embedding space. Self-supervised InfoNCE consistency and soft $\mathrm{SE}(3)$ invariance further encourage viewpoint-robust descriptors. In Stage 2, the distilled representation is adapted to registration through correspondence learning, density-aware point-dropout augmentation, and end-to-end pose optimization. With a single checkpoint, CVSD-Reg generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference. On KITTI, nuScenes, and HeLiPR, CVSD-Reg achieves strict success rate (SR@0.5\,m/$1^\circ$) of 97.7$\%$, 99.0$\%$, and 99.3$\%$, respectively, including 97.3$\%$ on sparse 16-beam Velodyne scans. It outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.

View source

Similar papers

2026

Z-REG: Domain-Robust Zero-Shot Point Cloud Registration via Equivariance-Aware Stabilization

Point cloud registration is crucial in real-world applications such as robotics and automation. While deep learning-based methods have made substantial progress in this domain, their generalization remains limited in heterogeneous environments. This limitation primarily arises from their strong reliance on domain-speci...

Jing-Yu Zhou, Yun-Feng Ma, Shuai Jiang et al. · 0 citations
Open access Sep 2026

SDA-Reg: large-scale dynamic scene point cloud registration with semantic dual-stream attention

Point cloud registration is a fundamental task in 3D vision, however, large-scale outdoor LiDAR point clouds, characterized by their immense size and structural complexity, present significant challenges in highly dynamic environments. Existing methods often employ semantic segmentation as preprocessing, offering the p...

Sheng-Jie Fu, Qi-Peng Cai, Zhao-Yuan Yao et al. · 0 citations
Preprint Sep 2026

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...

Yu-Hang Han, Youngseok Jang, Seungwon Roh et al. · 0 citations
Preprint Sep 2026

Unsupervised Point Cloud Registration via Training-Time Semantic Guidance

Unsupervised registration of large-scale LiDAR point clouds remains challenging due to the geometric ambiguity inherent in outdoor scenes, which degrades pseudo-label quality and leads to suboptimal convergence, particularly for sparse, low-resolution scans such as those from nuScenes. We reveal that registration model...

Kezheng Xiong, Shi-Yun Xu, Sheng Ao et al. · 0 citations
Preprint Aug 2026

GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

GhostPoint is proposed, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation, and introduces a predictor-level supervision scheme on sampled voxels from generated neighborhoods.

Mohamed Abdelsamad, Bin Yang, Michael Ulrich et al. · 1 citation
Conference Aug 2026

A unified geometry-aware framework for calibration-free cross-modal image-to-point cloud registration

Multi-view point cloud registration is critical for autonomous driving and 3D reconstruction. While incorporating image data enhances performance, current cross-modal methods relying on explicit fusion or geometric projection suffer from modality distribution gaps, calibration sensitivity, and high computational costs....

Yulin Hou, Yu Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.