Skip to content
Preprint

AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

AirAlign is proposed, a framework for RGB-only image-pair relative pose alignment for UAVs, using a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs.

Abstract

Unmanned aerial vehicle (UAV) navigation in modern low-altitude environments requires more accurate pose alignment in the final approach stage for target information acquisition or manipulation, making"last-meter"navigation increasingly important. However, severe viewpoint and appearance variations make this task challenging. To tackle this problem, we propose AirAlign, a framework for RGB-only image-pair relative pose alignment for UAVs. AirAlign uses a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs. In addition, to better utilize the limited training data, we split the training set into multiple scene-disjoint folds for unseen cross-validation and model selection. During inference, the predictions of the selected models are averaged to form the ensemble output of the overall framework. Experiments on the PairUAV challenge at the ACMMM 2026 Workshop on UAVs in Multimedia demonstrate the effectiveness and robustness of our method, while comprehensive ablation studies validate the contribution of each component.

View source

Similar papers

Open access Jul 2026

Monocular ORB-SLAM3 Evaluation for Multi-Altitude VTOL UAV Mapping

Abstract. Reliable visual localization is essential for long-range VTOL UAV mapping in GNSS-degraded environments. This paper presents a quantitative evaluation framework for monocular ORB-SLAM3 using a 66.48 km multi-altitude UAV mission and aerial-triangulation-derived camera poses as reference data. The workflow ass...

Ming-Jyun Yang, J. Jhan, Runmeng Tang · 1 citation
Preprint Aug 2026

DECO: Depth-Guided Co-Visibility Reasoning for Low-Altitude UAV Visual Localization

DECO is proposed, a DEpth-guided CO-visibility reasoning framework for low-altitude UAV visual localization that retains keypoints that are both visually distinctive and geometrically co-visible, improving feature matching and PnP-based pose estimation.

Yi-Bin Ye, Xichao Teng, Shuo Chen et al. · 0 citations
Jul 2026

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degre...

Xin Li, Siyuan Duan, Shang Wang et al. · 1 citation
Preprint Aug 2026

GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. Aft...

Jing-Pu Yang, Deming Tang, Yi-Lin Sun et al. · 1 citation
Open access Jul 2026

UAV Visual Localization in GNSS-Denied Environments

Abstract. Navigating Unmanned Aerial Vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments requires reliable autonomous localization techniques. This study proposes a vision-based localization framework utilizing satellite true orthophotos and Digital Surface Models (DSMs) as absolute geospat...

Tai-Cyuan Wang, Lai-Han Tsou, J. Jhan et al. · 0 citations
Open access Jul 2026

Real-Time Temporally Consistent Monocular 6D UAV Pose Estimation for Onboard Aerial Perception

AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related in...

Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.