Skip to content

Similar papers

Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access Jul 2026

MoReSP: A Multiobjective Mobility- and Reliability-Aware Scheduling Model for RSU-Assisted Vehicular IoT Networks

Simulation results show that MoReSP achieves the lowest admission-adjusted system cost across all evaluated scenarios, and demonstrate that MoReSP provides a reliable and balanced scheduling solution for dynamic V-IoT environments.

Muhammad Faisal Siddiqui, Adeel Iqbal · 0 citations
Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
Preprint Aug 2026

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application

This paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent and the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues.

Marcos Carvalho, Fatih Temiz, Shavbo Salehi et al. · 0 citations
Conference Jul 2026

Stochastically Enhanced Multi-Agent Federated Reinforcement Learning for Energy-Efficient and Quality-of-Service-Aware IoT Routing

This work addresses the challenges of energy efficiency and Quality of Service (QoS) in dynamic Internet of Things (IoT) networks, where traditional routing protocols fail to adapt to changing conditions. To overcome these limitations, a Stochastically Enhanced Multi-Agent Federated Reinforcement Learning (SE-MA-FRL) framework is proposed. The method integrates multi-agent reinforcement learning with federated learning and stochastic policy exploration to enable distributed, privacy-preserving, and adaptive routing decisions. The model uses parameters such as residual energy, link quality, queue length, and distance to optimize routing. Simulation results demonstrate that the proposed approach achieves a Packet Delivery Ratio (PDR) of 96.96%, outperforming conventional and existing RL-based methods, while also reducing delay, packet loss, and energy consumption. In conclusion, SE-MA-FRL provides a scalable, reliable, and energy-efficient routing solution suitable for largescale and dynamic IoT environments with strict QoS requirements.

P. R, Debasis De, Sheetal Belaldavar · 0 citations
Open access Jul 2026

Constraint-Aware Resource Exploration for Multi-Agent Collaborative Offloading in Mobile Edge Computing

A constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE) that achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.

Yuxuan Yang, Hexing Wang, Yang Zhou · 0 citations