The first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception is introduced, and ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, is developed and deployed on a physical UAV platform.
Abstract
We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC), explicitly connecting high-level task understanding, behavior planning, and low-level control. The datasets are collected from both real-world and simulated outdoor environments for training and evaluation. We further develop ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, and deploy it on a physical UAV platform. Experiments with representative VLMs and VLA models show that current UAV agents still struggle with behavior planning, viewpoint adjustment, and robust task completion in active perception. These results establish ActiveFly-Bench as a new testbed for embodied aerial intelligence.
By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Shenghong Yi, Lin Zhang, Muzian Li et al.· 0 citations
Deployable autonomy remains a key challenge for unmanned aerial vehicles (UAVs) operating in open-ended missions. Large language models (LLMs) and their multimodal variants, which can process visual and other sensory inputs, have introduced new capabilities for semantic perception, task reasoning, and language-conditio...
Ting-Quan Xiong, Jianning Zhan, Qiu-Wei Deng et al.· Drones· 0 citations
This work introduces MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments, and shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning.
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt et al.· arXiv.org· 0 citations
LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Shao-An Wang, Ao-Cheng Luo, Fei Huang et al.· 2 citations
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on imp...
Ting Huang, Yue Huang, Ze-Yu Zhang et al.· 0 citations
UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error...
Hao-Yu Zhang, Shuoxun Zhang, Peng Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.