GALoc, a geometry-first framework that replaces depth prediction with gravity-aligned wireframes that satisfy verticality and coplanarity by construction, is proposed and evaluated end-to-end on Structured3D, with calibrated noise on Gibson, and on real-world author-collected sequences.
Abstract
Floorplans are compact, appearance-invariant maps ideal for indoor localization, yet existing methods rely on depth networks that are brittle in cluttered scenes. We propose GALoc, a geometry-first framework that replaces depth prediction with gravity-aligned wireframes that satisfy verticality and coplanarity by construction. Given monocular RGB, camera intrinsics, relative poses, and IMU orientation, GALoc constructs a linear constraint matrix encoding verticality and coplanarity, and finds the camera gauge minimizing its smallest singular value via global search. The rectified wireframes are projected into bird's-eye-view layouts through a closed-form, FOV-consistent transformation and matched against the floorplan via metric-free SE(2) search. We evaluate end-to-end on Structured3D, with calibrated noise on Gibson, and on real-world author-collected sequences. When sufficient wall geometry is visible, GALoc matches or outperforms depth-based baselines -- achieving 88% sequential localization success at 0.1m over 100-step sequences on Gibson vs the baseline's 68% -- while abstaining in structure-blind scenes.
Monocular Gaussian SLAM must recover camera motion, surface structure, and appearance from an RGB sequence without metric depth input or benchmark geometry during reconstruction. This setting is challenged by scale-ambiguous predictions, spatially varying reliability, and the tendency of an unconstrained Gaussian map t...
Reliable 3D spatial understanding is essential for autonomous navigation, obstacle avoidance, and scene reconstruction. While state-of-the-art learned depth estimation techniques achieve high accuracy in-distribution, they often generalize poorly to novel viewpoints and altitudes. This paper presents a geometrically de...
Diksha Aggarwal, Rutvik Dagadkhair, Sanjana Srivastava et al.· 0 citations
Indoor building construction sites are demanding environments for visual SLAM, where variable lighting and repetitive, low-textured structures make the system drift over long trajectories, though structural elements such as walls remain distinguishable despite these conditions. These buildings are constructed according...
Asier Bikandi-Noya, Miguel Fernández-Cortizas, Muhammad Shaheer et al.· 0 citations
This work presents a different approach inspired by a human navigation technique called resection, that can perform direct ground to satellite image matching and localization without relying on external depth foundation models, and achieves faster, memory-efficient inference.
Hyeongsik Kim, Mincheol Kim, Heejoon Moon et al.· 0 citations
Transparent and specular surfaces pose a serious challenge to LiDAR-based SLAM and navigation because laser returns may pass through glass, leaving collision boundaries absent from the map. Prior work attempts to reconstruct the missing surfaces, but inaccurate obstacle placement can create the opposite failure: contam...
Han-Wen Guo, Zheng-Zhi Lin, Yu-Sen Xie et al.· 0 citations
Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.
Jia-Rong Lian, Zhen-Hua Xiao, Zhao-Yang Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.