Aug 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 52 references
Medicine
TL;DR
DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (DO-Conv).
Abstract
Accurate and robust vehicle localization is essential for autonomous driving. However, existing visual pose estimation methods often struggle in scenarios dominated by repetitive structures or sparse textures. These conditions lead to ambiguous predictions of 3D scene coordinates and a high proportion of structured outliers—erroneous predictions forming coherent clusters that deceive standard estimators. To address these limitations, this paper proposes DOU-Pose (Depthwise Over-parameterized U-shaped Pose estimation), a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline. The core idea is to enhance the discriminative capability of scene coordinate regression through improved feature extraction. Specifically, we replace standard convolutional layers with Depthwise Over-parameterized Convolution (DO-Conv), which introduces auxiliary learnable depthwise kernels during training to enrich the representational capacity of the network, while allowing their fusion into a single kernel for inference. Furthermore, a U-shaped regression network with transposed convolutions is designed to preserve spatial details and strengthen fine-grained geometric reasoning. The entire pipeline is trained end-to-end by coupling dense scene coordinate prediction with a differentiable robust estimator. Extensive experiments demonstrate that DOU-Pose achieves competitive performance on public benchmarks and clear robustness improvements on the self-collected Campus-AV dataset, especially in repetitive and low-texture outdoor driving scenarios.
LoFG is presented, a localization-oriented Feature Gaussian representation that unifies both stages within a single Gaussian scene and improves the robustness of sparse initialization and the accuracy of dense refinement, demonstrating potential for localization applications in AR, robotics, and visual navigation syste...
Zhen-Dong Xiao, Zi-Ling Wen, Jun Yin et al.· Multimedia Systems· 0 citations
Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirement...
Eulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno et al.· 0 citations
A scalable visual localization pipeline that combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation is introduced, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches.
Jonas Meyer, S. Nebiker, P. Theiler et al.· arXiv.org· 0 citations
This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.
Yasmin Loeper, Markus Gerke, P. Fanta-Jende· The International Archives o...· 0 citations
In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early warning, and safe autonomous driving. However, existing methods suffer from high sensitivity to camera extrinsics; even slight deviations (whether manifesting as transie...
Junsheng Du, Zhaocheng He, Yuhuan Lu· arXiv.org· 0 citations
PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.
Leon Jungemeyer, A. Magaña, Gautham Mohan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.