2026· IEEE Transactions on Network and Service Management· Vol 23, pp. 6721-6734· 0 citations· 35 references
Abstract
Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.
This paper investigates the problem of cooperative multiple unmanned aerial vehicles (UAVs) data collection for Internet of Things (IoT) networks in dense urban environments. Unlike existing studies that predominantly rely on idealized spatial models and average-based probabilistic channel models, this work explicitly accounts for realistic 3-D building distributions and deterministically models ground-to-air (G2A) channel blockages. We formulate a joint optimization problem to minimize the total task completion time, subject to stringent system throughput, flight dynamics, and energy constraints. To tackle the highly coupled challenges of node scheduling and trajectory planning, we propose a lightweight two-stage heuristic strategy for dynamic access control, along with a multi-agent reinforcement learning for trajectory planning. Crucially, to overcome the severe sparse-reward bottleneck inherent in complex 3-D obstacle avoidance, we introduce a Pheromone-based Reward Shaping (PRS) mechanism. By mathematically integrating the UAV’s kinematic state with deterministic environmental feedback, PRS effectively transforms the sparse-reward navigation challenge into a dense and smooth gradient, thereby profoundly accelerating policy convergence. Extensive simulations demonstrate that the proposed MATD3-PRS framework significantly outperforms representative baselines, achieving superior performance in task completion time, flight trajectory efficiency, and overall energy saving.
Haitao Chen, Xinfeng Deng, Zhe Wang et al.· IEEE Transactions on Cogniti...· 0 citations
A multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets and outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.
Shuanli Jia, Naiming Qi, Zheng Li et al.· Drones· 0 citations
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.
Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al.· 0 citations
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications, such as collaborative search and rescue and environmental monitoring. In multi-UAV cooperative tracking, accurate arrival-time coordination is important for improving collaborative task execution, but it remains challenging because of continuous action spaces, target maneuvering, uncertain time-to-go estimation, and inefficient exploration in multi-agent reinforcement learning. Specifically, a multi-UAV cooperative guidance environment is formulated, and the problem is modeled as a Markov decision process. To address the challenges of large action spaces and poor convergence in multi-agent reinforcement learning, an experience-guided MAPPO framework is introduced to enhance training efficiency and policy stability. Different from standard MAPPO, the proposed E-MAPPO introduces proportional-navigation-guided experience only during the early training stage to guide exploration, while the final policy is still optimized through the MAPPO objective. Subsequently, a composite reward function is designed by integrating distance-based heuristic terms with auxiliary guidance signals, thereby improving exploration efficiency and facilitating coordinated rendezvous and tracking of dynamic references. Comparative simulations with cooperative proportional navigation guidance (CPNG), sliding mode control (SMC), and standard MAPPO are conducted under different target motion scenarios. The results show that E-MAPPO reduces the average convergence step by 17.07% compared with MAPPO. In the straight-moving target scenario, E-MAPPO reduces the cooperative time error by 55.10% compared with CPNG and by 8.33% compared with MAPPO. In the S-type maneuvering target scenario, E-MAPPO reduces the cooperative time error by 55.81% compared with CPNG and by 9.52% compared with MAPPO. Monte Carlo experiments further verify its effectiveness and robustness. Additional robustness tests under Gaussian measurement noise, observation bias, and communication delay show that the proposed method maintains acceptable tracking accuracy and cooperative timing performance under different uncertainty conditions. In addition, the results indicate that the proposed method generalizes well to different types of maneuvering targets.
Hao Xiong, Minghu Tan, Xiaoyu Liu et al.· Drones· 0 citations