Skip to content

Multi-UAV Trajectory Planning for Dynamic Target Search: An LLM-Enhanced Multi-Agent Reinforcement Learning Algorithm

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 10294-10310 · 0 citations · 32 references

Abstract

Deploying Uncrewed Aerial Vehicles (UAVs) for dynamic target search in disaster response scenarios can reduce losses. This paper investigates multi-UAV cooperative trajectory planning for dynamic target search in a three-dimensional environment with static obstacles, aiming to maximize the number of searched targets and minimize the average uncertainty of the search area, while ensuring collision avoidance between UAVs and obstacles. Existing Multi-Agent Reinforcement Learning (MARL) based methods face the sparse reward problem in dynamic target search, which hinders planning feasible multi-UAV trajectories. Notably, Large Language Models (LLMs), with extensive pre-trained knowledge and powerful semantic reasoning capabilities, exhibit potential for designing high-quality reward functions to alleviate the sparse reward problem. Therefore, we propose an LLM-guided Multi-Agent Proximal Policy Optimization (LLM-MAPPO) algorithm, which leverages LLMs’ reasoning capabilities to guide MARL policy learning and plans multi-UAV trajectories for efficient dynamic target search. Specifically, we design an offline LLM reward shaping scheme that generates dense reward signals to mitigate the sparse reward problem. Moreover, we propose a dual-mode pheromone-based search mechanism to guide UAVs to respond promptly to changes in target positions. Experimental results demonstrate that LLM-MAPPO significantly outperforms compared algorithms in terms of the number of searched targets and average area uncertainty, while successfully avoiding collisions. In particular, LLM-MAPPO reduces the target search time by 71.4%.

View source

Similar papers

2026

Collaborative Trajectory and Resource Optimization in Multi-UAV MEC Under Jamming: An LLM-Guided MARL Framework

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.

Yeguang Qin, Jie Tang, Fengxiao Tang et al. · 0 citations
Open access Jul 2026

A Multi-UAV Cooperative Mission Planning Method Based on Multi-Agent Guided Soft Actor–Critic

A multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets and outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.

Shuanli Jia, Naiming Qi, Zheng Li et al. · 0 citations
Open access Jul 2026

An Experience-Guided MAPPO Framework for Multi-UAV Cooperative Tracking in Continuous Action Spaces

A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications, such as collaborative search and rescue and environmental monitoring. In multi-UAV cooperative tracking, accurate arrival-time coordination is important for improving collaborative task execution, but it remains challenging because of continuous action spaces, target maneuvering, uncertain time-to-go estimation, and inefficient exploration in multi-agent reinforcement learning. Specifically, a multi-UAV cooperative guidance environment is formulated, and the problem is modeled as a Markov decision process. To address the challenges of large action spaces and poor convergence in multi-agent reinforcement learning, an experience-guided MAPPO framework is introduced to enhance training efficiency and policy stability. Different from standard MAPPO, the proposed E-MAPPO introduces proportional-navigation-guided experience only during the early training stage to guide exploration, while the final policy is still optimized through the MAPPO objective. Subsequently, a composite reward function is designed by integrating distance-based heuristic terms with auxiliary guidance signals, thereby improving exploration efficiency and facilitating coordinated rendezvous and tracking of dynamic references. Comparative simulations with cooperative proportional navigation guidance (CPNG), sliding mode control (SMC), and standard MAPPO are conducted under different target motion scenarios. The results show that E-MAPPO reduces the average convergence step by 17.07% compared with MAPPO. In the straight-moving target scenario, E-MAPPO reduces the cooperative time error by 55.10% compared with CPNG and by 8.33% compared with MAPPO. In the S-type maneuvering target scenario, E-MAPPO reduces the cooperative time error by 55.81% compared with CPNG and by 9.52% compared with MAPPO. Monte Carlo experiments further verify its effectiveness and robustness. Additional robustness tests under Gaussian measurement noise, observation bias, and communication delay show that the proposed method maintains acceptable tracking accuracy and cooperative timing performance under different uncertainty conditions. In addition, the results indicate that the proposed method generalizes well to different types of maneuvering targets.

Hao Xiong, Minghu Tan, Xiaoyu Liu et al. · 0 citations
Open access Jul 2026

MULTI-UAV COORDINATED PATH PLANNING USING A MULTI-AGENT SOFT ACTOR-CRITIC ALGORITHM

An efficient way to resolve the curse of dimensionality, improve obstacle avoidance and cooperative formation control of UAVs was found and shows great prospects of practical application in such domains as military operations, search and rescue missions, transport automation and disaster management.

Qadir Talibov · 0 citations
Open access Aug 2026

A Trajectory Planning Method for UAVs in Dynamic Multi-Threat Environments Based on a Dynamic Multi-Objective Crow Search Algorithm

Trajectory planning, which determines a route from a starting position to a target position within a given airspace, is critical to unmanned aerial vehicle (UAV) mission execution. Many existing meta-heuristic approaches to three-dimensional (3D) trajectory planning aggregate competing requirements into a weighted cost and may suffer from limited adaptability when the environment changes. This paper formulates 3D UAV trajectory planning in dynamic multi-threat environments as a dynamic bi-objective optimization problem and proposes a multi-swarm dynamic multi-objective crow search algorithm (MDMCSA). The proposed method organizes objective-oriented swarms within a cooperative search framework and facilitates information exchange through archive sharing, thereby coordinating the search process among different objectives. The memory-time and diverse behavior strategies adjust search behaviors and solution perturbation to balance convergence and diversity. A hybrid change response strategy combines historical information reuse with diversity restoration after dynamic changes. Comparative experiments on dynamic benchmark problems and UAV trajectory planning scenarios demonstrate competitive convergence and adaptation performance, together with a favorable trade-off between solution quality and computational cost. Incremental ablation and parameter-sensitivity analyses further indicate the cumulative benefit of the integrated design and the stable performance of the selected parameter configuration across the tested settings.

Gengsong Li, Yi Liu, Qibin Zheng et al. · 0 citations
Open access Jul 2026

Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO

Low-altitude urban environments pose significant challenges to multi-UAV trajectory planning because of dense buildings, constrained airspace, dynamic obstacles, inter-UAV conflicts, and terminal-area congestion. This study proposes a hierarchical three-dimensional cooperative trajectory-planning framework integrating dynamic-risk-aware Ant Colony Optimization (ACO) with cooperative Model Predictive Control–Gray Wolf Optimizer (MPC-GWO). Environmental costs and predicted dynamic-obstacle risks are incorporated into the ACO global search to generate risk-aware reference trajectories, while a sliding-window GWO improves trajectory smoothness and execution feasibility. During online execution, cooperative MPC-GWO combines dynamic-obstacle prediction, inter-UAV separation constraints, reconfigurable formation switching, and goal-neighborhood safety control to achieve adaptive obstacle avoidance, cooperative replanning, and orderly terminal arrival. Thirty-run Monte Carlo simulations show that the proposed method achieves a success rate of 93.3% ± 25.4% and the highest composite score of 96.20 ± 5.30, with zero dynamic-obstacle and inter-UAV collisions. Ablation experiments verify the effectiveness of the dynamic prediction, formation reconfiguration, and terminal safety-control mechanisms. The average online replanning time remains below 0.5 s, demonstrating satisfactory safety, coordination, adaptability, and real-time performance in small- to medium-scale simulated urban scenarios.

Yuhan Wang, Pengfei Zhang, Yawen Li et al. · 0 citations