Physics-Guided Learning for Monocular Visual Object Localization in Indoor Environments
Abstract
Accurate object localization is essential for enabling autonomous operation of indoor robotic systems. As a low-cost, compact, and flexibly deployable solution, monocular visual object localization (MVOL) is highly applicable to lightweight embedded robotic platforms. However, conventional data-driven MVOL methods suffer from inherent limitations of 2D-to-3D ill-posed mapping, which is caused by the incapability of constraining the spatial physical logic of real scenes, resulting in severe depth ambiguity, inaccurate scale estimation, and physically unreasonable predictions. To address these issues, this paper proposes a novel physics-guided monocular visual localization framework termed PC-IMVL for indoor scenarios. The PC-IMVL integrates deep visual perception with embedded physical modeling, which explicitly introduces spatial physical constraints into the network optimization process and builds a physical consistency-aware loss function to regularize 3D position and pose estimation. Combined with a lightweight tailored architecture, the framework enables efficient and reliable embedded deployment. Offline experiments and real-world online tests validate the effectiveness of the proposed method. PC-IMVL yields average absolute errors (AE) of 0.095–0.333 m, reducing the localization error of early fusion methods by more than 50%. Within a working distance of 3–4 m, it achieves a relative error (RE) of 2.4% and a horizontal viewing angle error (VAE) below 2°, outperforming existing state-of-the-art MVOL methods. The effectiveness of the physical guidance mechanism is verified. This work provides a practical high-precision localization solution for embedded indoor robotic systems.