Skip to content
Open access

Hierarchical Graph-Attention Multi-Agent Reinforcement Learning for Safe-Separation-and-Collision-Avoidance Coordination of Heterogeneous UAV Swarms

Jul 2026 · Drones · Vol 10, pp. 508 · 0 citations · 52 references

TL;DR

Results suggest that HG-MARL is a promising simulation-validated framework for civilian UAV swarm coordination in collision-and-separation-critical and communication-degraded environments.

Abstract

Safe-separation-and-collision-avoidance unmanned aerial vehicle (UAV) swarms are increasingly used for inspection, emergency response, environmental monitoring, and search-and-rescue support in cluttered airspace where communication links may be delayed, degraded, or intermittently unavailable. These applications require heterogeneous vehicles to maintain situational awareness, allocate tasks, and avoid hazards under partial observability and changing team topology. To address these challenges, this paper proposes a Hierarchical Graph-Attention Multi-Agent Reinforcement Learning architecture (HG-MARL) for safe-separation-and-collision-avoidance heterogeneous UAV swarm coordination. The proposed framework decomposes the task into high-level resource allocation and low-level local-control execution, uses graph attention for changing swarm topology, and applies Transformer memory, action masking, potential-field reward shaping, and domain-randomized simulation training. In the multi-scenario simulation summaries, HG-MARL achieves 92.9%, 89.8%, and 82.6% task success in Scenarios A–C, respectively, improving upon MAPPO by 15.1, 21.4, and 20.1 percentage points. Summary-statistic Welch tests show that all six HG-MARL comparisons against MAPPO and QMIX yield p<0.01 with large effect sizes. Fair-control, reward-sensitivity, communication-degradation, safety-ablation, training-stability, latency, and transfer-oriented stress tests further support the contributions of the integrated architecture. The validation scope is simulator-based, with platform-level flight/HIL evaluation discussed as future work. These results suggest that HG-MARL is a promising simulation-validated framework for civilian UAV swarm coordination in collision-and-separation-critical and communication-degraded environments.

Read PDF

Similar papers

Open access Aug 2026

Collision-aware cooperative multi-UAV path planning with hierarchical PPO-LSTM

The results indicate that separating waypoint-level strategy from recurrent local execution improves mission reliability and collision avoidance in the tested grid environments, while larger random-map benchmarks, fully controlled MAPPO/QMIX comparisons, and continuous 3-D simulation remain important future work.

Alparslan Güzey · 0 citations
Aug 2026

FALCON-MASAC: Formation-Aware Attention-Enhanced Leader-Guided Control-Barrier Optimization for Safe Multi-UAV Formation Navigation in Dynamic 3-D Environments

FALCON-MASAC is presented, a safety-integrated multi-agent reinforcement learning framework that decomposes this task into four complementary layers: a hierarchical leader-follower paradigm that pairs a pre-trained virtual leader with followers learning a distributed cooperative policy, and a bypass-side commitment coo...

Yi-Ming Shang, Chang-Ping Du, Rui Yang et al. · 0 citations

Temporal coordination aware reinforcement learning for multi-agent UAV navigation in dynamic environments

T-CARE is introduced, a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations and achieves 100% success, 0% collision rate, and no observed persistent starvation or deadlock.

Abhudaya Shrivastava, Z. Obradovic · 0 citations
Review Jul 2026

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

A multi-agent deep reinforcement learning framework that addresses issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation is proposed, demonstrating strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenar...

Yuhuang Su, Nabil Aouf · 0 citations
Open access Aug 2026

Modeling Dynamic Obstacle Avoidance Strategy of Drone Swarms Combined with Multi-Agent Reinforcement Learning

The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.

X.-H. Fang, K. Chen, Cheng-Hao Ren et al. · 0 citations
Conference Aug 2026

Attention-based MADDPG with dual-buffer experience replay for cooperative multi-UAV target tracking

An Attention-based Multi-Agent Deep Deterministic Policy Gradient algorithm was developed for cooperative multi-unmanned aerial vehicle target tracking in dynamic environments. The study addressed information redundancy and association weight allocation between individual unmanned aerial vehicles and the swarm during c...

Qing-Lin Han, Hongmei Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.