Skip to content
Review Open access

Fusion-Oriented Deep Learning-Enhanced Visual SLAM: A Review

2026 · Computers, Materials & Continua · Vol 89, pp. 1-10 · 0 citations · 93 references

TL;DR

This paper presents a systematic review of deep learning-enhanced VSLAM, with a particular focus on how learning models are incorporated into classical simultaneous localization and mapping (SLAM) pipelines and how they function within the overall system.

Abstract

: Visual simultaneous localization and mapping (VSLAM) is a key technology for mobile robotics, autonomous driving, and embodied intelligence, enabling self-localization, environment reconstruction, and scene understanding. Although conventional geometric methods have achieved notable success, their performance often degrades in challenging conditions, such as low-texture scenes, severe illumination changes, dynamic interference, and long-term environmental variations. Recent advances in deep learning have created new opportunities to improve VSLAM through stronger feature representations, learned priors, semantic perception, and emerging map representations. At the same time, the increasing adoption of learning-based modules has raised important questions about integration strategies, generalization, interpretability, and real-time deployment. This paper presents a systematic review of deep learning-enhanced VSLAM, with a particular focus on how learning models are incorporated into classical simultaneous localization and mapping (SLAM) pipelines and how they function within the overall system. To provide a unified perspective, existing methods are organized into five categories according to their fusion interfaces with geometric SLAM pipelines: observation-level interfaces, constraint/prior/weight-level interfaces, solver-level interfaces, representation-level interfaces, and system-level integration interfaces. Based on this taxonomy, representative approaches are comparatively analyzed for accuracy, robustness, efficiency, and deployability. In addition, this review summarizes common design principles, including geometric consistency constraints, error propagation characteristics, and typical failure modes, and further discusses open challenges and future directions such as lightweight deployment, cross-domain adaptation, dynamic map modeling, and long-term consistency maintenance. This review aims to provide a structured reference for the analysis, design, and deployment of learning-enhanced VSLAM systems.

Read PDF

Similar papers

Review Open access Aug 2026

Review of Visual SLAM: From Geometry to Learning with Challenges and Future Directions

In recent decades, visual SLAM has evolved enormously from geometric approaches to learning-based approaches. There are various algorithms that are used for a variety of scenarios with different complexity levels. Each of these algorithms has its own strengths and limitations. It becomes difficult to select a visual SL...

M. Kamble, D. Watvisave, Omkar Watvisave et al. · 0 citations
Review Sep 2026

Monocular Depth Estimation from a Single Image: Progress and Opportunities

Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative f...

Mu-Xin Liu, Xiaoyang Lyu, Yang-Tian Sun et al. · 0 citations
Open access Sep 2026

Deep Learning-Based Visual Semantic Perception Technology for Mobile Vehicles

With the large-scale implementation of scenarios such as logistics warehousing and park inspections, the demand for autonomous environmental perception by mobile robots continues to grow. Low-cost pure vision semantic perception has become a core technology for ensuring autonomous safe navigation of these robots. Tradi...

Jia-Wei Sun · 0 citations
Conference Aug 2026

Monocular distance estimation: from geometric foundations and deep learning innovations to industrial deployment challenges

It is concluded that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion and self-supervised frameworks.

Zi-Kang Fan, Zi-Hao Xiang, Jiang-Sheng Liu · 0 citations
Aug 2026

Universal Representation for Real-World Misaligned Infrared-Visible Image Fusion.

Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification o...

Jin-Yuan Liu, Zengxi Zhang, Jiahao Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.