X-GeoP2P, a coarse-to-fine localization pipeline that combines adapted versions of GAReT and X-VGGT, is presented, a coarse-to-fine localization pipeline that combines adapted versions of GAReT and X-VGGT.
Pixel-level cross-view geo-registration aims to align a query image (e.g., drone) to a geo-referenced satellite map so that every query pixel can be mapped to real-world GPS coordinates. Despite strong progress in cross-view geo-localization, existing benchmarks largely provide only GPS labels, limiting evaluation to a single coordinate per image and leaving dense geodetic alignment underexplored. We introduce SkyReg, a dataset and standardized benchmark for pixel-level drone-to-satellite geo-registration, providing dense per-pixel geo-location supervision across diverse settings (orthographic and perspective), scene types (urban, landmark-centric, suburban/rural), and camera configurations. Using SkyReg, we evaluate a broad set of baselines spanning retrieval, feature matching, homography-based alignment, and feed-forward 3D reconstruction. Finally, cross-view pairs from SkyReg, we train a geometry-aware reconstruction pipeline that achieves state-of-the-art results,improving performance by a significant margin.
Qingyang Liu, D. Shatwell, P. Kulkarni et al.· 0 citations
Cross-view geo-localization (CVGL) aims to match images captured from different viewpoints, such as drone and satellite imagery. Existing methods primarily focus on single-image matching, overlooking the potential of leveraging temporal information from drone image sequences. To address this, we propose Spatio-Temporal Alignment and Refinement (STAR), a novel framework that effectively leverages the spatio-temporal dependencies of drone sequences for effective matching with satellite images. Notably, this is among the first attempts to explore drone-sequence-based matching in CVGL, further improving retrieval performance through spatio-temporal consistency. Specifically, we design the Video Vision Transformer (ViViT) Sequence Alignment Module (VSAM) to effectively extract and align sequential features and introduce a Pseudo-Temporal Compensation (PTC) strategy to prevent the temporal modeling branch from degenerating when processing static satellite inputs, maintaining architectural symmetry across dynamic and static views. In addition, we design the Dynamic Background Partitioning Module (DBPM) to adaptively segment features and improve foreground–background distinction. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that our method significantly outperforms state-of-the-art approaches, highlighting the benefits of incorporating sequence information in CVGL.
Ziqian Mo, Yu-Xi Sun, Sen Jia et al.· IEEE Transactions on Geoscie...· 0 citations
Cross-view localization (CVL) estimates the pose of a ground image by matching it to a geo-referenced satellite image. To bridge the extreme viewpoint gap, mainstream pipelines rely on Bird's-Eye-View (BEV) transformations or 2D-to-3D lifting. However, deriving 3D structures from a single ground image is fundamentally ill-posed, causing these methods to endure geometric distortions and computational costs during 3D lifting or BEV projection. Furthermore, relying on external depth foundation models to resolve this introduces latency and remains susceptible to noisy predictions. In this work, we present a different approach inspired by a human navigation technique called resection, that can perform direct ground to satellite image matching and localization without relying on external depth foundation models. The key insights of our method are that (i) ground keypoints can be translated into azimuthal rays on the satellite map, and (ii) these rays ideally converge at the user location. Exploiting this geometric constraint through direct line-to-point correspondences, we introduce a minimal Azimuthal Ray Convergence (ARC) solver to identify the intersection, alongside an ARC loss to optimize the matching network. By eliminating dependencies on computationally heavy BEV transformations and external depth foundation models, our approach achieves faster, memory-efficient inference, while its explicit feature matching ensures straightforward compatibility with existing frameworks. Experiments on VIGOR and KITTI demonstrate that ARC-Loc maintains competitive localization accuracy compared to recent approaches, highlighting its practicality.
Hyeongsik Kim, Mincheol Kim, Heejoon Moon et al.· 0 citations
Experiments with a state-of-the-art CVL model on OpenCVL show that incorporating noisy in-the-wild data consistently improves performance on clean test sets, suggesting a promising direction for scaling CVL with diverse real-world imagery.
Zimin Xia, Mubariz Zaffar, Junfan Fu et al.· 0 citations
In global navigation satellite system (GNSS)-denied scenarios, cross-view geolocalization (CVGL) provides an effective solution for autonomous uncrewed aerial vehicle (UAV) localization. However, viewpoint and scale differences across platforms may cause variations in texture structures and spatial distributions for the same geographic region, making it more difficult to learn consistent and discriminative cross-view representations in CVGL. Meanwhile, visually similar but geographically distinct samples can reduce the separability between true matches and hard negatives. To address these issues, we propose the frequency-spatial feature enhancement and dynamic margin constraint (FSDC) network, which integrates the frequency-aware recalibrated spatial (FARS) module and the margin-based dynamic contrastive learning (MDCL) strategy. The FARS module enhances local structural representations through bidirectional complementary interaction between frequency-domain and spatial structural features, guiding the shared encoder to learn more consistent cross-view representations, while the MDCL strategy imposes adaptive bounded constraints on hard-negative samples to improve feature discriminability and training stability. Experimental results on three benchmarks show that FSDC provides a favorable tradeoff among cross-view retrieval accuracy, model complexity, and inference efficiency.
Wei-Quan Wang, Yan-Fei Peng, Lei Ma et al.· IEEE Geoscience and Remote S...· 0 citations
Abstract. In the initial response to wildfires, securing rapid and accurate geographic information is essential. However, helicopter imagery acquired on-site often lacks precise sensor metadata, such as camera pose and internal parameters, making the application of georeferencing difficult. In particular, obliquely captured wildfire imagery presents additional registration challenges due to severe viewpoint changes, scale variations, and low-texture environments. This study proposes an automated georeferencing pipeline capable of operating under these constraints. The proposed method consists of five stages: preprocessing, image retrieval, feature extraction and matching, Exterior Orientation Parameters (EOP) estimation, and orthomosaic generation. An initial Area of Interest (AOI) is defined using inaccurate initial position data, and the Region of Interest (ROI) within the reference map is obtained through a ResNet50-based image retrieval approach. Subsequently, virtual Ground Control Points (GCPs) are generated through deep learning-based feature matching. Elevation data is then assigned using a Digital Elevation Model (DEM), and EOP are estimated via Perspective-n-Point (PnP) and RANSAC algorithms. Intermediate frames are initialized via interpolation and refined through bundle adjustment to produce the final orthomosaic. Experimental results demonstrated that utilizing SuperGlue and LightGlue complementarily increased the number of successfully georeferenced intervals from 5 to 9. Furthermore, a minimum RMSE of 28.30 m was achieved in the most accurate interval. This method proves that by automating the feature-based georeferencing process, practical geographic information can be rapidly provided for initial disaster response, even in sensor-limited environments.
Seongyun Kim, Jeonghyo Oh, J. Cheon et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.