GCAFormer-VO: A Geometry and Correspondence-Aware Transformer for Visual Odometry
Abstract
Visual odometry (VO) estimates camera motion from image sequences and is essential for robotics, autonomous driving, and AR/VR. Robust VO remains challenging because large viewpoint changes and strong parallax make reliable cross-frame motion cues difficult to capture, especially in the presence of visual disturbances such as occlusions and motion blur. Consequently, this article proposes a geometry- and correspondence-aware Transformer-based VO (GCAFormer-VO) that improves motion estimation by explicitly injecting camera geometry cues and strengthening local cross-frame interaction. First, we introduce geometry-aware patch embedding (GA-PE), which augments per-frame feature grids with spatial encoding and bearing-ray cues derived from camera intrinsics. Building on these geometry-augmented features, we propose a cross-frame parallax module that captures motion evidence through two parallel pathways. Difference cross-frame window attention (DCF-WA) performs efficient windowed cross-frame attention within a short temporal radius, and a parallax alignment encoder (PAE) extracts parallax-sensitive pairwise cues using feature differences, local correlation, and edge responses. Their outputs are integrated by a lightweight parallax fusion module (PFM) to form a unified pairwise representation. Finally, a decoupled pose head regresses rotation and translation using task-specific spatial attention pooling. Experimental results on the KITTI and TartanAir datasets demonstrate that GCAFormer-VO achieves improved accuracy and robustness compared with strong baselines, and qualitative analyses further show stable motion-estimation behavior in representative challenging cases.