This paper proposes AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback and boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
Abstract
Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fail...
Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al.· IEEE Robotics and Automation...· 0 citations
Quadrotor UAVs are increasingly deployed in complex missions that demand reliable autonomous navigation and robust obstacle avoidance. Traditional modular pipelines suffer from cumulative latency, motivating a shift toward end-to-end learning-based methods. However, this paradigm still faces two fundamental challenges....
Yan-Jie Liu, Teng-Da Yang, Zi-Han Li et al.· IEEE Robotics and Automation...· 0 citations
CoNav-UAV is proposed, which explicitly models the target-oriented vision-and-language navigation task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone.
Jun-Ru Song, Wenhao Zhang, Yang Yang et al.· 2 citations
We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments. To bridge the gap between abstract semantics and low-level control, we employ a parallelized ensemble of six behavior-conditioned Model Predictive Path Inte...
Hanbing Zhang, Fang-Guo Zhao, Ze-Rui Li et al.· 0 citations
With the rapid advancement of autonomous flight technology, there is an increasing demand for higher precision and advanced navigation techniques, with permissible distance errors often restricted to a few meters. Furthermore, the ubiquitous deployment of Unmanned Aerial Vehicles (UAVs) necessitates the integration of...
Nga Vu Quynh, P. N. Huu, Thanh Han-Trong· IEEE International Conferenc...· 0 citations
Vision-based Unmanned Aerial Vehicles (UAVs) often suffer from navigation failures in dead ends due to limited sensing accuracy and range. To address this challenge, this paper proposes a systematic solution for efficient dead-end prediction and avoidance. The proposed method introduces a lightweight neural network to...
Rui-Bin Zhang, Lun Pan, Zelong Xia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.