Skip to content

An Intelligent DRL‐Based Framework for Reliable UAV Swarm Communications in Dynamic Environments

Jul 2026 · Expert systems · Vol 43 · 0 citations · 29 references

TL;DR

Results show that topology‐aware cooperative learning can provide a scalable and practical solution for reliable UAV communication in future intelligent aerial and 6G‐enabled networks, particularly in dense deployments.

Abstract

Unmanned Aerial Vehicle (UAV) networks are increasingly deployed in dynamic environments where reliable and low‐latency communication is critical. However, high mobility, intermittent connectivity, spectrum limitations, and energy constraints make conventional static communication protocols inadequate for maintaining stable, dependable links. To address these challenges, this paper proposes TRC‐MAPPO, a topology‐aware reliability‐constrained multi‐agent deep reinforcement learning framework for adaptive UAV swarm communication. The routing problem is formulated as a constrained decision‐making task that jointly considers packet delivery reliability, end‐to‐end delay, link stability, bandwidth usage, and energy consumption. Unlike single‐agent DRL baselines, TRC‐MAPPO represents the swarm as a dynamic communication graph and combines local relay selection with centralized training, enabling cooperative routing decisions under time‐varying network conditions. Simulations are conducted in a controlled, dynamic UAV environment, with the same mobility and traffic settings used for all methods. The proposed framework is compared with DQN and PPO over different swarm densities. Results show that TRC‐MAPPO achieves a higher packet delivery ratio, lower end‐to‐end delay, and more efficient energy behaviour, particularly in dense deployments. These findings indicate that topology‐aware cooperative learning can provide a scalable and practical solution for reliable UAV communication in future intelligent aerial and 6G‐enabled networks.

View source

Similar papers

Open access 2026

QEGT-Based Adaptive Routing for Energy-Efficient and Reliable Communication in UAV Swarm Networks

This study proposes an intelligent Q-learning-enhanced Evolutionary Game Theory (QEGT) routing mechanism for USNs that leverages game-theoretic incentives and Q-learning to adaptively select strategies.

Anita Murmu, Saurabh Kumar Srivastava, Nuthan Chingeetham et al. · 0 citations
Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
Conference Jul 2026

Joint Trajectory and Spectrum Optimization for Anti-Jamming UAV Swarms: A DRL Approach

Reliable link maintenance is currently a critical bottleneck for unmanned aerial vehicle (UAV) swarm communications in complex electromagnetic environments where UAVs encounter both external malicious jamming and internal interference. Most recent studies have treated trajectory design and resource scheduling as decoupled problems or employed standard deep reinforcement learning methods to handle static spectral scenarios. However, these approaches lead to frequent link breakages and slow convergence when dealing with dynamic topologies and spatiotemporal interference. To tackle this challenge, we proposes a joint spatial-spectral adaptive coordination (JSSAC) framework and a deep recurrent attentionbased Q-network (DARQN) approach, utilizing a multi-head attention mechanism to intelligently aggregate heterogeneous neighbor features, thereby enhancing the swarm's adaptability to dynamic network topology. Moreover, considering that the spatial distribution of drones fundamentally determines the upper bound of the signal quality, we designed a communicationaware potential field mechanism that incorporates real-time signal-to-interference-plus-noise ratio feedback. Simulation results demonstrate that compared to DQN and DRQN algorithms, the proposed algorithm achieves transmission success rates of over 92%, representing improvements of 17% and 8% respectively, while also accelerating convergence speed.

Miao Liu, Nan Qi, Hua Jiang et al. · 0 citations
Open access 2026

Multi-Tier UAV Swarm Deployment for JRC Systems: Distributed Optimization With Learning-Based Adaptation

Simulation results confirm the effectiveness of distributed optimization and DRL-based coordination for scalable, resilient, and adaptable UAV deployment in disaster response and other mission-critical scenarios.

A. Abdellatif, Amr E. Aboeleneen, Mohamed M. Abdallah et al. · 0 citations
#edge computing Open access Aug 2026

Collaborative resource allocation in UAV-assisted MEC networks: A heterogeneous MAPPO scheme

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) is a key enabler for meeting the stringent low-latency and energy-efficiency requirements of emerging low-altitude economy applications. However, achieving these objectives remains challenging due to dynamic environments, limited communication and computation resources, and the heterogeneity of network entities. This paper investigates the long-term joint optimization framework that minimizes system-wide latency and energy consumption simultaneously by coordinating UAV association, subchannel selection, uplink/downlink power allocation, and computational resource distribution. This sequential decision-making process is formulated into a partially observable Markov decision process (POMDP) to account for localized observations and dynamic channel states. To solve it, we propose a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices (UDs) and UAVs act as heterogeneous agents. This architecture utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers. Numerical results demonstrate that the proposed scheme effectively navigates the high-dimensional action space and achieves superior convergence and cost reduction compared to benchmarks, including PPO, independent PPO (iPPO), Q-learning multi-agent extension (QMIX), value decomposition networks (VDN), independent deep Q-network (iDQN), and genetic algorithm (GA).

Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al. · 0 citations
Review Open access Jul 2026

Development of UAV Swarm Ad-Hoc Network Communication Technology for Emergency Scenarios: A Review

Major disasters such as earthquakes, floods, and wildfires can rapidly destroy terrestrial communication infrastructure, producing an extreme operating environment in which power, road, and network outages compound one another. Owing to their rapid deployability, flexible networking, and three-dimensional mobility, unmanned aerial vehicle (UAV) swarms are being studied as a flexible component of emergency communication systems. This paper reviews UAV swarm ad-hoc network communication technology for emergency scenarios. It examines the technical characteristics and applicability boundaries of three network architectures---flat, hierarchical clustering, and space-air-ground integrated---and surveys recent advances in routing and medium access, intelligent networking optimization, and transmission and security assurance. Particular attention is given to the reported performance and applicability of emerging approaches, including reinforcement-learning-based adaptive routing, decentralized federated learning, digital twins, and semantic communication, under highly dynamic and resource-constrained conditions. Drawing on studies of emergency routing, post-disaster data collection, semantic forwarding, and multi-layer coverage, the paper assesses current validation methods and outlines research directions in energy use, scalability, security, resilience, and standardization. Its contribution is a cross-layer comparison that relates architecture choices to protocol requirements, implementation costs, and validation maturity.

Yihang Ren, Huatao Zhu, Jie Zhang · 0 citations