TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation
TOLiD is proposed, a self-supervised pretraining method for LiDAR representation learning that addresses the gap between dense ViT token representations and sparse 3D encoders by coupling a LiDAR backbone with a student Vision Transformer initialized from a frozen VFM teacher and applying supervision over compatible pa...