Skip to content

Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach

Jul 2026 · arXiv.org · Vol abs/2607.18604 · 0 citations · 16 references
Computer Science Engineering

TL;DR

Simulation results demonstrate that the proposed Hierarchical LLM-driven control framework significantly reduces collision rates and improves aggregate system throughput compared to existing baselines.

Abstract

The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kinematics and multi-tier network handovers. While Deep Reinforcement Learning (DRL) offers rapid tactical control, it lacks the zero-shot strategic reasoning required to quickly adapt to dynamic Integrated Terrestrial and Non-Terrestrial Networks (ITNTNs). Conversely, Large Language Models (LLMs) excel at semantic reasoning but suffer from high inference latency, rendering them unsuitable for real-time aerodynamic control. To bridge this gap, we propose a novel Hierarchical LLM-driven control framework. A massive cloud-based LLM deployed on a High-Altitude Platform Station (HAPS) manages slow-timescale global load balancing, while lightweight edge-LLMs on individual UAVs translate local observations into tactical sub-goals. These sub-goals guide a fast-timescale physical DRL controller to execute collision-free, handover-aware trajectories. Simulation results demonstrate that our agentic architecture significantly reduces collision rates and improves aggregate system throughput compared to existing baselines.

View source

Similar papers

2026

Hallucination-Aware Hierarchical LLM for Autonomous UAV Mobility Control: A Safe Reinforcement Learning Approach

This paper proposes SafeGPT, a hierarchical framework that integrates generative pretrained transformer (GPT)-based large language models (LLM) with safe reinforcement learning (safe-RL). SafeGPT addresses the large-scale random multi-point tour problem (RMPT) for multiple unmanned aerial vehicles (UAVs). The target platforms are large electric vertical take-off and landing (eVTOL) class rotary-wing UAVs. Such platforms suit wide-area logistics and infrastructure inspection rather than small drones. The hierarchical architecture enables scalable UAV coordination by centralizing strategic planning and distributing local route computation, mitigating the complexity issues of conventional methods. To mitigate LLM hallucination risks, i.e., the generation of plausible routes that violate physical constraints, a safe-RL method is implemented as a constrained Markov decision process (MDP) with Lagrangian actor-critic optimization, which promotes constraints on energy consumption, communication efficiency, and duplicate visits. Comprehensive simulations demonstrate SafeGPT’s performance advantages in battery efficiency, communication reliability, and travel distance across problem scales from 500 to 2000 waypoints. Even in demanding scenarios, constraint violation rates remain below specified thresholds. These results establish SafeGPT as an effective solution for large-scale UAV route planning, respecting operational constraints without compromising efficiency.

Hyojun Ahn, Seungcheol Oh, Gyusun Kim et al. · 0 citations
Sep 2026

A Multi-UAV Cooperative Navigation Method Based on Policy Decomposition Structure

Cooperative navigation of multiple unmanned aerial vehicles (UAVs) in disaster search-and-rescue scenarios is challenging due to dense obstacles, partial observability, and strong inter-agent coupling, which often result in path conflicts, collision risks, and limited policy generalization. To address these challenges, this paper proposes a Multi-Agent Deep Deterministic Policy Gradient framework with a Graph-Attention-based Staged Actor (GS-MADDPG). Under a centralized training and decentralized execution paradigm, GNNs are employed to model local interaction relationships among UAVs, enabling effective information aggregation and cooperative decision-making under partial observability. Furthermore, the Actor network is decomposed into perception, goal-guidance, and feature fusion subnetworks, allowing hierarchical decoupling and coordinated integration of local obstacle avoidance behaviors and global navigation objectives. Simulation results conducted in a complex three-dimensional urban environment demonstrate that, compared to traditional methods, GS-MADDPG improves the navigation success rate, robustness, and generalization performance. When the obstacle density reaches 50% and the number of UAVs increases from 2 to 10, the navigation success rate of GS-MADDPG is approximately 40% higher than that of the benchmark algorithm; even in cases with higher obstacle density, GS-MADDPG still achieves a relatively high success rate. This verifies its effectiveness in multi-UAV cooperative navigation for search and rescue tasks.

Li Tan, Hai-Xia Zhao, Jia-Qin Chai et al. · 0 citations
Conference Aug 2026

LLM-Guided Task Planning with A*-Based Navigation for Autonomous UAVs

Dynamic unmanned aerial vehicle (UAV) missions require online adaptation not only to geometric changes but also to evolving task-level constraints, such as energy depletion, facility unavailability, and newly imposed visits. This paper presents an event-driven hierarchical planning framework that integrates large language models (LLMs) with deterministic A*-based navigation. Upon detecting a mission event, an LLM revises only the remaining sequence of symbolic facility visits, while a layered verifier independently checks response structure, task semantics, energy and precedence constraints, and geometric executability. Each accepted task sequence is subsequently converted into collision-free flight segments by A*, thereby isolating language-based task reasoning from safety-critical path execution. Twelve LLMs are evaluated on a common urban map under four controlled conditions: recharge-triggered recovery, temporary no-fly-zone activation, mandatory waypoint insertion, and a joint event combining energy, visitation, and airspace constraints. An exhaustive task-sequence oracle provides a scenario-specific lower bound for evaluating route optimality. Across 48 model–scenario trials, 33 produced mission-valid plans, with the best-performing models reaching or closely approaching the oracle route length in all individual-event scenarios. Performance declined markedly under the joint event, where only 41.7% of the models generated valid plans, highlighting the difficulty of satisfying interacting mission constraints.

Lu-Wei Liao, Hong-Zong Li, Li Zhang et al. · 0 citations
Open access Aug 2026

Optimizing 3D UAV navigation via a high-efficiency hybrid RRT*-DQN approach

This paper develops a hybrid path-planning method, rapidly-exploring random tree star (RRT*)-deep Q-network (DQN), which combines the fast global search capability of RRT* with the deep learning-based heuristic prediction of a DQN.

Abhishek Bajpai, A. Abhinav, N. Tiwari · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations
Preprint Aug 2026

CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning

CoNav-UAV is proposed, which explicitly models the target-oriented vision-and-language navigation task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone.

Junru Song, Wenhao Zhang, Yang Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.