Monocular Visual-Inertial Odometry via Explicit 3D Motion Modeling
Abstract
Visual-inertial odometry (VIO) estimates motion by fusing camera and inertial data. In learning-based monocular VIO, depth, scale, and motion are tightly coupled, while IMU cues are difficult to impose as geometric constraints on implicit 2D features. This paper presents a pose estimation method based on explicit 3D motion modeling and geometric enhancement. Depth and optical flow form dense 3D motion features that encode scene structure and inter-frame motion. A plane-sweep cost volume improves multi-view consistency, while short-term priors from IMU preintegration guide visual- inertial fusion. On KITTI, the method obtains average translational and rotational errors of 1.95% and 0.81°/100 m, respectively. Compared with the best baseline, UnVIO, these errors are reduced by 13.5% and 6.5%, respectively. The results validate the proposed components and show that explicit 3D geometry and inertial priors improve pose estimation. This geometry-aware design supports reliable localization for robots, unmanned aerial vehicles, and autonomous driving in visually degraded scenes.