Review of Visual SLAM: From Geometry to Learning with Challenges and Future Directions
Abstract
In recent decades, visual SLAM has evolved enormously from geometric approaches to learning-based approaches. There are various algorithms that are used for a variety of scenarios with different complexity levels. Each of these algorithms has its own strengths and limitations. It becomes difficult to select a visual SLAM algorithm for a particular scenario. This article provides the evolution of front-end architecture of visual SLAM from pure geometric methods to learned features, deep optical flow, semantic segmentation, object-level SLAM and differentiable rendering. It then provides the advancement in backend from filtering techniques to optimization methods, enhancing loop closure, outlier methods, fusion of sensors and learning enhanced backends. This article also showcases the strengths, weaknesses, and use cases of algorithms and their comparison across varied environmental conditions and computational requirements. The paper provides information about the trade-off between some performance parameters along with a few selection guidelines. Towards the end, this article provides major challenges in the implementation of visual SLAM along with research areas such as dynamic scenes, adverse illumination, long-term map management, scalable dense mapping, uncertainty estimation, sensor calibration, semantic consistency, benchmarking standards, and deployment on resource-limited platforms. Continuous evolution of visual SLAM algorithms helps better perception and understandability of the environment, making them more accurate and robust.