Aug 2026· International Conference on Laser, Optics and Optoelectronic Technology· Vol 14314, pp. 143143H - 143143H-6· 0 citations· 21 references
Engineering
TL;DR
It is concluded that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion and self-supervised frameworks.
Abstract
Monocular distance estimation, a fundamental yet ill-posed problem in computer vision, has evolved from rigid geometric constraints to flexible deep learning paradigms. This review provides a comprehensive analysis of this transition, categorizing methodologies into traditional geometric-based, supervised, and self-supervised frameworks. We examine how traditional methods (e.g., size priors and ground-plane constraints) offer interpretability but struggle with environmental robustness. In the deep learning era, we detail the shift from CNN receptive field limitations to Transformer based global dependency modeling (e.g., DPT, MonoViT) and the mathematical progression from continuous regression to adaptive depth binning. A significant focus is placed on self-supervised mechanisms, specifically the "photometric consistency" assumption and landmark innovations like SfM-Learner’s joint optimization and Monodepth2’s auto masking strategy. Finally, we synthesize performance benchmarks on the KITTI and NYU Depth V2 datasets to highlight current bottlenecks—namely, scale ambiguity and domain shift. The review concludes that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion.
Monocular Depth Estimation (MDE) is one of the most rigorously studied problems in modern computer vision, yet it is fundamentally ill-posed. Recovering absolute three-dimensional geometry from a single two-dimensional projection is mathematically impossible without strong inductive priors. This paper presents a struct...
Hasan Mahmud Shanto, Mohammad Tofiqul Islam, Muhammad Ryan Hasan et al.· AIUB Journal of Science and...· 0 citations
Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative f...
Mu-Xin Liu, Xiaoyang Lyu, Yang-Tian Sun et al.· 0 citations
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp a...
Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al.· 0 citations
Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Pr...
XiDepth, a lightweight architecture based on the XiNet operator block, designed to enhance feature extraction while maintaining low computational complexity and energy demand, is proposed.
Elena Izzo, Riccardo Toniolo, Lamberto Ballan· 1 citation
Self-supervised monocular depth estimation (MDE) eliminates the reliance on expensive ground-truth depth annotations and has emerged as a powerful approach for a wide range of vision applications. However, current lightweight networks are hampered by two critical challenges: the limited representation capacity of stati...
Dong-Liang Wang, Ming Jin, Xiang-Qian Fang et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.