Skip to content
Preprint

Scalix: Uncertainty-Aware Scale-Consistent Monocular SLAM

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

Scalix is proposed, a real-time monocular SLAM framework that achieves metric-scale state estimation by integrating learned depth cues into a probabilistic factor-graph formulation, leading to improved scale consistency through multi-view data associations.

Abstract

Cameras are ubiquitous sensors in robotics due to their compact form factor and the perceptual richness captured through visual information. Monocular SLAM enables robots to understand the environment with a minimum setup, however, it inherently suffers from scale ambiguity. A common solution is to provide multi-modal sensor configurations, such as visual-inertial systems, where scale is observable unless the robot navigates under a constant-velocity motion, a common scenario in mobile robotics. With the advent of deep-learning, geometric foundation models have been used to address this problem, but the depths maps are often noisy and scale-inconsistent across frames. In this paper, we propose Scalix, a real-time monocular SLAM framework that achieves metric-scale state estimation by integrating learned depth cues into a probabilistic factor-graph formulation. By augmenting existing monocular depth models with both per-pixel depth uncertainty and per-frame scale uncertainty, Scalix treats scale predictions as independent measurements within its optimization, leading to improved scale consistency through multi-view data associations. Experiments in large-scale outdoor and indoor environments demonstrate state-of-the-art performance on both metric and up-to-scale benchmarks while maintaining real-time operation and generalization.

View source

Similar papers

Preprint Oct 2026

CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

Autonomous vehicles often suffer from limited perception due to occlusions, blind spots, limited sensor range, and the complex nature of surrounding environments. Multi-agent collaborative perception (CP) addresses these challenges by allowing vehicles to share sensory information and reconstruct the scene cooperativel...

Soham Pahari, Sudip Das, Arindam Das et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DepthWorld: 3D World Model for Robot Manipulation

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-fr...

Jai Bardhan, Josef Sivic, Vladimír Petrík · 0 citations
Preprint Oct 2026

Sensor-Layout-Agnostic Navigation via Geometric Observation Canonicalization

Existing visual navigation policies are inherently bound to fixed camera configurations, creating a fundamental barrier to zero-shot deployment across heterogeneous robot sensor layouts. To overcome this limitation, we present an embodiment-informed navigation policy capable of generalizing across diverse depth sensor...

Welf Rehberg, Kostas Alexis · 0 citations
Preprint Sep 2026

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

A camera and an IMU are the minimal sensor setup for metric localization and dense mapping, yet classical visual--inertial filters must wait for parallax before they start and then retain only sparse landmarks. Feed-forward geometry models, in contrast, predict dense structure from a few images but provide neither metr...

J. Mahmoud, Arthur Movsesyan, Mikhail Iumanov et al. · 0 citations
Preprint Aug 2026

Beyond Relative Geometry: Metric-Aware Geometry Perception for Robotics

Recent embodied models increasingly leverage geometric representations to improve spatial reasoning and robotic manipulation. However, existing reconstruction methods only reconstruct relative geometry with arbitrary scales, causing predicted object dimensions and spatial distances to vary across scenes, viewpoints, an...

Fengjun Zhong, Cong-Jia Chen, Zhao Liu et al. · 0 citations
Preprint Aug 2026

VeloBins: Learning Velocity and Its Uncertainty via Bins and Error-Conditioned Gaussian Labels for Aerial Inertial Odometry

Inertial odometry (IO) is critical for aerial robots, where aggressive maneuvers and poor lighting degrade visual sensors. Recent learning-based IO methods improve traditional integration-based approaches by learning motion priors from IMU and platform-specific sensors, then fusing the predictions within an extended Ka...

Maulana Bisyir Azhari, Seungwook Lee, Dong-Hun Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.