Skip to content
Conference

Deep State-Space Monocular Visual Odometry with Learnable Kalman Filtering

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · pp. 3155-3158 · 0 citations

TL;DR

A monocular visual odometry method that combines deep temporal features with a Kalman filtering module based on a state-space model, which improves pose estimation robustness and adaptability to incomplete observations and verifies the effectiveness of combining classical filtering theory with deep learning for monocular visual odometry.

Abstract

Monocular visual odometry is important in autonomous driving, robotics, and related fields, and has attracted increasing attention in computer vision.Traditional geometric methods and end-to-end deep learning methods have achieved promising results in monocular visual odometry, but they still face limitations in temporal consistency, uncertainty representation, and abnormal observation handling.To address these issues, this paper proposes a monocular visual odometry method that combines deep temporal features with a Kalman filtering module based on a state-space model. By explicitly modeling the temporal evolution of motion states in the state space, the method introduces continuity constraints into the pose estimation process. At the same time, the state transition matrix, as well as the process noise covariance matrix and the measurement noise covariance matrix, are learned by neural networks, which enables the system to adaptively adjust the fusion weight between prediction and observation. Experimental results on the KITTI VO dataset show that, compared with the baseline method, the proposed model improves pose estimation robustness and adaptability to incomplete observations, which verifies the effectiveness of combining classical filtering theory with deep learning for monocular visual odometry.

View source

Similar papers

Conference Aug 2026

Monocular Visual-Inertial Odometry via Explicit 3D Motion Modeling

Visual-inertial odometry (VIO) estimates motion by fusing camera and inertial data. In learning-based monocular VIO, depth, scale, and motion are tightly coupled, while IMU cues are difficult to impose as geometric constraints on implicit 2D features. This paper presents a pose estimation method based on explicit 3D mo...

Qian Li, Fan Bai · 0 citations
Preprint Sep 2026

FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry

Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate camera motion and 3D structure in one pass, pose estimation over long videos remains challenged by computational cost, long-context ambiguity, and temporal instability. To a...

Meng-Li Shih, Shih-Yang Su, Yu-Liang Zou et al. · 0 citations
Preprint Sep 2026

SFVO: Decoupled Confidence-Guided Stereo-Flow Visual Odometry with Bidirectional PnP

Deep learning-based visual odometry (VO) has achieved significant progress, yet most existing methods focus on a monocular approach, which suffers from scale ambiguity. Stereo VO provides real metric by its nature, but remains less studied in deep learning VO due to its high computational cost and modeling complexity....

Kai Zhang, Guo-Yang Zhao, Jun Ma · 0 citations
Review Sep 2026

Monocular Depth Estimation from a Single Image: Progress and Opportunities

Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative f...

Mu-Xin Liu, Xiaoyang Lyu, Yang-Tian Sun et al. · 1 citation
Conference Open access 2026

Xvins: Boosting State Estimation Robustness via Hybrid Temporal Tracking and Efficient Deep Feature Extraction

This work proposes XVINS, a hybrid VIO frontend integrating XFeat—a lightweight deep feature extractor—into the optimization-based VINS-Fusion framework, presenting XVINS as a viable, real-time state estimation solution for agile Micro-Aerial Vehicles (MAVs) and mobile platforms.

Thura Peou, Sarot Srang, Lychek Keo · 0 citations
Open access Sep 2026

Estimación de pose por filtrado complementario de marcadores y odometría

Accurate estimation of rigid object pose is a fundamental requirement in vision-based robotic. In eye-in-hand configurations, estimates obtained through Perspective-n-Point (PnP) from visual markers exhibit frame-to-frame noise.  We propose a recursive method that combines visual observations of ArUco markers with the...

Víctor Quesada Conejero, F. Real, Sergio Garrido Jurado · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.