RAT-CVGL: Rank-Aware Transformer for Cross-View Geo-Localization
Cross-view geo-localization (CVGL) is a critical task that determines the geographic position of a query image via retrieving its corresponding counterparts across heterogeneous visual domains, such as ground-level, drone, and satellite views. Despite significant progress, recent existing approaches primarily focus on enhancing cross-view feature representations and often rely on post hoc, domain-specific (Sat. $\rightarrow $ Dro. or Sat. $\leftarrow $ Dro.) re-ranking strategies to refine retrieval results. In this work, in contrast, we propose a universal rank-aware transformer (RAT)-CVGL framework, a learning-based re-ranking approach that integrates a RAT module with a dual-stream training strategy. In particular, the RAT module leverages rotary positional embedding (RoPE) and self-attention mechanisms to effectively model the relative positional relationships among neighboring features across different domains. To further improve ranking consistency and generalization, we introduce a dual-stream training strategy, complemented by an auxiliary ranking loss and data augmentation techniques. Comprehensive experiments on the University-1652 demonstrate the efficacy of our RAT module, and the proposed RAT-CVGL can achieve superior retrieval performance over existing methods.