Skip to content

Hallucination-Aware Hierarchical LLM for Autonomous UAV Mobility Control: A Safe Reinforcement Learning Approach

2026 · IEEE Transactions on Networking · Vol 34, pp. 6827-6842 · 0 citations · 45 references

Abstract

This paper proposes SafeGPT, a hierarchical framework that integrates generative pretrained transformer (GPT)-based large language models (LLM) with safe reinforcement learning (safe-RL). SafeGPT addresses the large-scale random multi-point tour problem (RMPT) for multiple unmanned aerial vehicles (UAVs). The target platforms are large electric vertical take-off and landing (eVTOL) class rotary-wing UAVs. Such platforms suit wide-area logistics and infrastructure inspection rather than small drones. The hierarchical architecture enables scalable UAV coordination by centralizing strategic planning and distributing local route computation, mitigating the complexity issues of conventional methods. To mitigate LLM hallucination risks, i.e., the generation of plausible routes that violate physical constraints, a safe-RL method is implemented as a constrained Markov decision process (MDP) with Lagrangian actor-critic optimization, which promotes constraints on energy consumption, communication efficiency, and duplicate visits. Comprehensive simulations demonstrate SafeGPT’s performance advantages in battery efficiency, communication reliability, and travel distance across problem scales from 500 to 2000 waypoints. Even in demanding scenarios, constraint violation rates remain below specified thresholds. These results establish SafeGPT as an effective solution for large-scale UAV route planning, respecting operational constraints without compromising efficiency.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.