DAP-Pose accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment, and incorporates physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift.
Abstract
Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.
SAPose is proposed, a lightweight geometry-guided pose-estimation framework that integrates multi-task keypoint prediction with geometric pose recovery and achieves competitive pose estimation accuracy with a compact model size and efficient inference.
DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (D...
Xin'an Qiu, Li-Wen Wang, Zezheng Dong et al.· Italian National Conference...· 0 citations
GS-CPE (Gaussian Splatting based Camera Pose Estimation), a coarse-to-fine framework for 6-DoF camera pose estimation that unifies geometry-based coarse pose estimation with robust 3D Gaussian Splatting based pose refinement, is introduced.
Real-time 6DoF object pose estimation on resource-constrained hardware remains challenging, as accurate correspondence-based and refinement pipelines typically rely on non-differentiable PnP/RANSAC stages or costly iterative refinement, while recent foundation-model-based approaches incur inference costs that are prohi...
P. Kühn, Duc Anh Nguyen, Saptarshi Neil Sinha et al.· 0 citations
Sen-Cap is proposed, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera that achieves state-of-the-art performance on major metrics on Human-M3 and FreeMotion, as well as strong cross-domain performance on LiDARHuman26M and RELI11D.
Aoru Xue, Yu-Jing Sun, Yiming Ren et al.· 0 citations
AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related in...
Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al.· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.