Skip to content

Quantum-Enhanced Multi-Agent Reinforcement Learning for Ubiquitous LLM Inference via Embodied UAV Swarms

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 11006-11019 · 1 citation · 50 references

Abstract

6G mobile edge networks are emerging as a key infrastructure for ubiquitous large language model (LLM) inference services. However, conventional edge routing to nearby or well-connected servers falls short for efficient edge LLM inference, as it may miss the user’s KV cache and trigger costly prefill recomputation. To address this challenge, this paper studies an edge inference system assisted by an embodied UAV agent swarm, where UAVs actively sense user mobility and neighboring UAV states to make local decisions on trajectory control, user association, and inference-request routing. The goal is to improve KV-cache reuse while maintaining reliable wireless connectivity, thereby maximizing the system effective token throughput under energy and QoS constraints. We then formulate the joint optimization as a mixed-integer non-linear program and further cast the sequential UAV decision-making process as a decentralized partially observable Markov decision process. To obtain scalable decentralized policies under partial observations, we propose Q-MAA2C, a quantum-enhanced multi-agent advantage actor-critic algorithm for embodied UAV swarm control and inference routing. Q-MAA2C uses quantum actors for local action selection and an entangled split critic for swarm-level value estimation, enabling coordinated policies from partial observations with reduced raw observation exchange. Simulation results indicate that Q-MAA2C yields comparable reinforcement learning rewards to the fully classical baseline while reducing the number of convergence episodes by about 43%. Additionally, the proposed method enhances the system effective token throughput by up to about 134% over other competing methods.

View source

Similar papers

Conference Jul 2026

Joint AoI and SWIPT-Aware Scheduling via Multi- Agent Deep Reinforcement Learning

This work investigates the joint optimization of Age of Information (AoI) and energy harvesting (EH) in wireless edge computing systems, where edge servers not only process IoT data but also act as wireless power suppliers via simultaneous wireless information and power transfer (SWIPT). Building upon the asynchronous model-free fractional multi-agent reinforcement learning framework and the Lyapunov drift-plus-penalty (DPP) concept, we design a fractional-based reward function for AoI and construct a virtual queue to enforce long-term energy stability under battery storage constraints. The overall reward is formulated as a weighted sum, capturing the trade-off between timeliness and energy sustainability, with update decisions, task offloading, and power splitting ratios as key control variables. Simulation results demonstrate that the developed multi-agent deep reinforcement learning approach achieves superior AoI–energy trade-offs compared to related baseline algorithms. These findings highlight the effectiveness of our framework in balancing information freshness and sustainable energy harvesting under resource-constrained edge environments.

Kuang-Ting Liu, Jain-Shing Liu, Wan-Ling Chang · 0 citations
Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
2026

Optimizing Information Freshness in Satellite-UAV IoRT Networks: A Heterogeneous Multi-Agent Approach

In satellite-UAV assisted communication networks, jointly optimizing the UAV’s trajectory and the multi-agent scheduling decisions to minimize the age of information (AoI) is a notoriously challenging problem. The complexity is compounded by the fundamental heterogeneity between the satellite and UAV agents, including their disparate action spaces, partial observations, and differing energy-consumption and communication-cost penalties. To address this, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose a novel heterogeneous multi-agent compound-action proximal policy optimization (HMACPPO) algorithm. HMACPPO leverages a centralized training with decentralized execution (CTDE) framework, using role-specific decentralized actors together with agent-specific centralized critics conditioned on the global state. Specifically, the UAV employs a compound PPO (CPPO) actor for its hybrid action space, while the satellite uses a PPO actor for discrete scheduling. Extensive simulations show that HMACPPO outperforms the compared baselines, and that the resulting coordinated policy effectively manages the trade-off between AoI, UAV energy consumption, and operational cost.

Weijie Zhou, Mengjie Yi, Yan Zhang et al. · 0 citations
2026

Multi-Agent Model-Based Reinforcement Learning for Decentralized Spectrum Sharing in Low-Altitude Economy

Rapid advances in drone technology, combined with the growing congestion of terrestrial transport networks, are driving the emergence of the low-altitude economy. Uncrewed Aerial Vehicles (UAVs) are increasingly deployed for low-altitude economy applications such as urban logistics and transportation, yet their expansion is constrained by the scarcity of spectrum resources. Although Multi-Agent Reinforcement Learning (MARL) offers a promising decentralized approach to improve spectral efficiency of UAVs, existing MARL methods suffer from high training costs, often requiring extensive environmental interactions. To overcome these limitations, we propose a novel Multi-Agent Model-Based reinforcement learning algorithm for decentralized spectrum sharing among UAVs in the low-altitude economy, which we denote as MAMBA-UAV. Adopting a Centralized Training with Decentralized Execution (CTDE) paradigm, MAMBA-UAV equips each UAV with a learned world model that captures compact environmental representations and predicts system dynamics. These world models are then utilized during MARL training to simulate interactions, thereby reducing the reliance on repeated real-environment rollouts. Through comprehensive simulations, we demonstrate that MAMBA-UAV substantially reduces the number of environmental interactions required for UAVs to achieve competitive spectrum-sharing performance, lowering training costs while maintaining high performance.

Tianle Li, Peixi Peng, Qingyu Liu et al. · 0 citations
Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
Preprint Aug 2026

Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks

The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connectivity. However, intermittent contacts, finite buffers, and limited message time-to-live (TTL) often give rise to sparse delivery and congestion, leading to substantial end-to-end performance degradation. To address this challenge, this study explores the joint optimization of decentralized opportunistic routing and controllable unmanned aerial vehicle (UAV) flight, aiming to enlarge future contacts through discrete UAV headings while enabling per-node replication under contact-limited observations. Building upon this architecture, we study cooperative factored routing--UAV control under centralized training and decentralized execution (CTDE) and propose JUROR (Joint UAV flight and Opportunistic Routing, based on the proximal policy optimization (PPO) framework. In our design, we first cast the problem as a factored partially observable Markov decision process with sequential motion--routing coupling and a per-step team reward; subsequently, decentralized actors act on local observations while a training-time critic uses global statistics, and an optional multi-horizon hotspot predictor provides auxiliary supervision. Simulation results over four traffic modes demonstrate effective gains over PRoPHET and MaxProp, while retaining contact-limited decentralized execution.

Xiao Wang, Shun Yang · 0 citations