Skip to content
Conference

LiDAR-Supervised Monocular Depth Estimation via Cross-Modal Supervision

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 927-932 · 0 citations · 15 references

Abstract

Accurate depth perception is a cornerstone of autonomous driving, yet LiDAR sensors—the primary source of metric depth—remain costly and operationally complex. In this paper, we propose a cross-modal supervision framework that uses sparse LiDAR depth maps solely during training, enabling camera-only dense depth inference at test time. A ConvNeXt-base encoder with an FPN neck and a lightweight depth decoding head is trained with a log-scale L1 loss, gradient consistency term applied exclusively at valid LiDAR pixels (~0.7% pixel density), and an image-guided edge-aware smoothness loss operating on all pixels, alongside a two-phase backbone freeze-then-unfreeze strategy to stabilize early convergence. Evaluated on a large-scale Korean highway dataset of 64,840 frames, our model achieves AbsRel of 0.0675, RMSE of 3.907 m, and $\delta \lt 1.25$ accuracy of 0.943, demonstrating that ultra-sparse LiDAR supervision is sufficient to train competitive monocular depth estimators.

View source

Similar papers

Preprint Sep 2026

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...

Yu-Hang Han, Youngseok Jang, Seungwon Roh et al. · 0 citations
Review Open access Jul 2026

The Evolution of Monocular Depth Estimation:From Spatial Regression to GenerativeFoundations and the Reliability Gaps

Monocular Depth Estimation (MDE) is one of the most rigorously studied problems in modern computer vision, yet it is fundamentally ill-posed. Recovering absolute three-dimensional geometry from a single two-dimensional projection is mathematically impossible without strong inductive priors. This paper presents a struct...

Hasan Mahmud Shanto, Mohammad Tofiqul Islam, Muhammad Ryan Hasan et al. · 0 citations
Open access Sep 2026

MACalib-Net: Spatiotemporal multi-attention cooperative network for LiDAR-camera extrinsic calibration

Extrinsic calibration accuracy is a critical bottleneck for LiDAR-camera fusion in autonomous driving. To address motion dynamics and cross-modal disparities, this paper proposes MACalib-Net, a spatiotemporal multi-attention cooperative network for LiDAR-camera extrinsic calibration. The proposed method introduces a...

Jian-Hui Li, Ya-Bin Ding, Qing-Po Xu · 0 citations
Preprint Aug 2026

Vernata: Self-Supervised Learning of LiDAR Point Representations

Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...

Oliver Lemke, Alexander Liniger, Abel Gawel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.