Skip to content
Open access

MA-HEAD-Net: Adaptive Rule-Guided Multi-Agent DRL for AoI Minimization in UAV-Assisted Emergency Networks

Aug 2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 9962-9978 · 0 citations · 65 references
Computer Science

Abstract

In post-disaster scenarios, uncrewed aerial vehicles (UAVs) are critical for establishing emergency communication networks. For time-critical rescue missions, information freshness is crucial, as decisions based on outdated data can lead to ineffective control actions. This paper investigates age of information (AoI) minimization for UAV-assisted emergency communications with heterogeneous emergency services. To characterize heterogeneous emergency services with bursty arrivals and different packet-size requirements, we model packet arrivals using a Markov-modulated Poisson process (MMPP) and adopt finite blocklength (FBL) theory to capture the coupling among transmission duration, packet completion, and AoI evolution. To balance delay-tolerant long-packet transmission and urgent short-packet response, we propose a mini-slot-embedded scheduling mechanism with adaptive checkpoint-interval selection. To solve the joint optimization problem of UAV trajectory control, user scheduling, and checkpoint-interval selection, we propose an adaptive rule-guided multi-agent deep reinforcement learning (MADRL) framework, named Multi-Agent Hybrid Expert-Algorithmic Decision Network (MA-HEAD-Net). MA-HEAD-Net incorporates communication-domain rule priors into a gated multi-head policy, where adaptive gates adjust the influence of rule-prior logits and learned policy logits for different sub-tasks. The policy and gating components are jointly optimized under the multi-agent proximal policy optimization (MAPPO) training framework, enabling rule-guided decision making while preserving the flexibility of policy learning. Simulation results show that MA-HEAD-Net improves policy-formation efficiency compared with representative MADRL baselines and achieves better AoI performance than both learning-based and heuristic baselines, demonstrating its effectiveness in dynamic UAV-assisted emergency communication scenarios.

Read PDF