This work introduces Universal Referring, a generalized UAV referring task that jointly expands the query modality and the output cardinality, and presents UAV-URNet, a detection-style baseline that maps heterogeneous queries into a shared query space and predicts variable-size target sets through set prediction.
Abstract
Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) benchmarks and methods, however, are largely built around text-only queries and single-object outputs, which limits their applicability to practical UAV scenarios involving reference images, multimodal instructions, absent targets, and multiple valid target instances. To address this gap, we introduce \emph{Universal Referring}, a generalized UAV referring task that jointly expands the query modality and the output cardinality. We construct \emph{UniRef-UAV}, a multimodal benchmark that supports text-only, image-only, and text+image queries with modality-dependent target cardinality, where text-only and text+image queries admit no-target, single-target, and multi-target grounding while image-only queries focus on existence-aware single-instance grounding. It also provides in-domain and cross-domain evaluation protocols for visual-query generalization. We further present \emph{UAV-URNet}, a detection-style baseline that maps heterogeneous queries into a shared query space and predicts variable-size target sets through set prediction. Extensive experiments show that UAV-URNet provides a stable and reproducible baseline with more consistent no-target discrimination and a more lightweight, reproducible implementation than large general-purpose MLLMs. Additional domain analysis, query-representation analysis, and ablation studies demonstrate that multimodal queries help reduce visual-query ambiguity and promote a more unified query--target alignment space. The annotations, visual query crops/images, train/validation/test splits, evaluation scripts, and baseline code will be made publicly available to facilitate reproducible research.
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from multi-view visual references. To study this underexplored problem, we introduce AGOS...
UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error...
Hao-Yu Zhang, Shuoxun Zhang, Peng Ye et al.· 0 citations
DECO is proposed, a DEpth-guided CO-visibility reasoning framework for low-altitude UAV visual localization that retains keypoints that are both visually distinctive and geometrically co-visible, improving feature matching and PnP-based pose estimation.
Yi-Bin Ye, Xichao Teng, Shuo Chen et al.· 0 citations
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, differences in acquisition time and imaging platform between UAV and reference imagery introduce substantial cross-domain appearance and viewpoint shifts, challenging robust six...
Xin Li, Si-Yuan Duan, Shang Wang et al.· Knowledge-Based Systems· 1 citation
This work proposes GrabVG, a novel visual grounding framework inspired by human visual search that generates a compact set of reliable object hypotheses through distillation-guided proposal induction and text-aware hypothesis filtering, substantially reducing background distractions and semantic mismatches.
Chaowei Wang, Yan Di, Jingjun Sun et al.· 0 citations
Key insights are provided for optimizing cross-lingual vision-language alignment of LVLMs and advancing practical multimodal applications in UAV scenarios and differences in language structures: East Asian languages depend on word order for semantic expression, whereas Western languages feature complex lexical morpholo...
Jue Chen, Penghui Huang, Ran Ding et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.