Skip to content
Preprint

Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

Aug 2026 · 0 citations · 10 references
Computer Science Engineering

TL;DR

Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios, and the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.

Abstract

Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. The optimization problem is NP-hard, due to the non-convex, non-smooth nature of the max-min fairness objective. To overcome these constraints, we formulate the BS placement as a Markov Decision Process (MDP) and systematically benchmark four DRL schemes: a discrete single-agent Deep Q-Network (DQN), a spatially partitioned Multi-Agent DQN, a continuous single-agent Deep Deterministic Policy Gradient (DDPG), and a geographically partitioned multi-agent DDPG framework. Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios. Additionally full coverage is achieved, and a fairness Jain's index of 0.94 is obtained. Finally, the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.

View source

Similar papers

Preprint Jul 2026

Optimal Base Station Placement for Beyond 5G Networks with Non-Convex Topology

This paper investigates the optimal placement of a millimeter-wave (mmWave) base station (BS) within a realistic U-shaped environment with non-convex topology. The problem is challenging and NP-hard due to the non-convex topology and the non-convex objective functions which are the sum-rate maximization and max-min fairness, the latter being additionally non-smooth. To address this challenge, the BS placement is formulated as a Markov Decision Process (MDP). Then, we propose two deep reinforcement learning (DRL) techniques: First, the deployment area is discretized into a grid and optimized using a Deep Q-Network (DQN). Second, the U-shaped region is partitioned into continuous subspaces, where a Deep Deterministic Policy Gradient (DDPG) agent is dedicated to each subspace then the best BS placement is selected among partitions. Results demonstrate that optimal placement achieves full coverage and yields a Jain index of 0.99. Furthermore, the proposed partitioned multi-space DDPG achieves better solution than DQN with lower complexity.

Mohamed M. H. Shalma, Amr Mansour, Ahmed El-Mahdy · 0 citations
2026

Dimension-Independent Multi-Agent DRL for Multi-Cell Interference Mitigation

Multi-agent deep reinforcement learning (DRL) offers a promising framework for inter-cell interference mitigation in multi-cell networks. In such networks, each cell is associated with an agent that learns from its local environment to maximize a reward, such as spectral efficiency. To effectively mitigate inter-cell interference, agents typically share model weights or local experiences with one another or with a central node. However, the exchange of such information incurs significant communication overhead in each communication round between the central node and the individual agents, posing a major bottleneck to efficient multi-agent DRL-based inter-cell interference mitigation. This paper presents a novel dimension-independent multi-agent DRL algorithm for multi-cell interference mitigation. By leveraging zeroth-order optimization, the proposed algorithm reduces the communication overhead from <inline-formula> <tex-math notation="LaTeX">$\mathcal {O}(d)$ </tex-math></inline-formula> to <inline-formula> <tex-math notation="LaTeX">$\mathcal {O}(1)$ </tex-math></inline-formula>, where <inline-formula> <tex-math notation="LaTeX">$d$ </tex-math></inline-formula> denotes the shared information dimension. This is achieved by exchanging only a constant number of scalar values between the central node and the agents in each communication round, independent of the dimension <inline-formula> <tex-math notation="LaTeX">$d$ </tex-math></inline-formula> of the shared weights or experiences. The proposed algorithm is evaluated on millimeter-wave networks with varying numbers of cells, demonstrating its effectiveness for interference mitigation. Specifically, under universal frequency reuse, the total sum-rate increases almost linearly with the number of cells. Simulation results show that the proposed algorithm effectively mitigates interference and maximizes spectral efficiency in line-of-sight (LoS), non-LoS, and mixed environments, while significantly reducing communication overhead.

M. Dahal, Mojtaba Vaezi · 0 citations
Conference Jul 2026

Joint AoI and SWIPT-Aware Scheduling via Multi- Agent Deep Reinforcement Learning

This work investigates the joint optimization of Age of Information (AoI) and energy harvesting (EH) in wireless edge computing systems, where edge servers not only process IoT data but also act as wireless power suppliers via simultaneous wireless information and power transfer (SWIPT). Building upon the asynchronous model-free fractional multi-agent reinforcement learning framework and the Lyapunov drift-plus-penalty (DPP) concept, we design a fractional-based reward function for AoI and construct a virtual queue to enforce long-term energy stability under battery storage constraints. The overall reward is formulated as a weighted sum, capturing the trade-off between timeliness and energy sustainability, with update decisions, task offloading, and power splitting ratios as key control variables. Simulation results demonstrate that the developed multi-agent deep reinforcement learning approach achieves superior AoI–energy trade-offs compared to related baseline algorithms. These findings highlight the effectiveness of our framework in balancing information freshness and sustainable energy harvesting under resource-constrained edge environments.

Kuang-Ting Liu, Jain-Shing Liu, Wan-Ling Chang · 0 citations
Conference Jul 2026

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

Yu-Heng An · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations
Open access Aug 2026

Multi-Objective Reinforcement Learning for Smart Planning of Electric Vehicle Charging Stations

The popularity of electric vehicles (EVs) is growing at a fast pace, creating a need for the strategic deployment of charging stations (CSs) to provide enough coverage, cost effectiveness, and compliance with grid and urban planning regulations. The deployment of large-scale infrastructure under multiple, often conflicting constraints remains a challenging engineering decision-making problem. In this paper, we propose a hybrid optimization framework that combines greedy initialization with reinforcement learning to efficiently explore the charging station deployment problem. The proposed approach employs Q-learning and Deep Q-Network (DQN) agents to iteratively refine the initial deployment while simultaneously optimizing deployment cost, charging demand coverage, and operational utility under practical planning constraints. The constraints include grid capacity limitations, renewable energy utilization, and fairness considerations. The proposed framework is evaluated in realistic urban scenarios. The experimental results demonstrate that the reinforcement learning (RL) approach achieves superior trade-offs among competing objectives compared to baseline heuristic strategies, while maintaining computational scalability for large candidate location sets. The proposed framework demonstrates stable performance across three evaluated deployment scenarios, indicating its potential applicability to increasingly complex charging infrastructure planning problems. The proposed methodology is scalable to other complex engineering planning and resource allocation problems characterized by multi-objective trade-offs and dynamic constraints. Beyond improving optimization performance, the proposed framework contributes to sustainable transportation planning by supporting the efficient deployment of electric vehicle charging infrastructure. Optimized charging station placement promotes greater accessibility to charging services, encourages electric vehicle adoption, reduces unnecessary travel associated with charging activities, and contributes to lower greenhouse gas emissions. Consequently, the proposed methodology provides decision-makers with a scalable and intelligent planning tool that supports the transition toward more sustainable and energy-efficient urban mobility systems.

A. Bousia · 0 citations