Skip to content
Open access

ShellMean-MAPPO: A Conflict-Aware MARL Framework for Downlink Resource Allocation in Multi-Shell LEO Satellite Networks

2026 · IEEE Open Journal of the Communications Society · Vol 7, pp. 10193-10213 · 0 citations · 55 references

TL;DR

This work proposes a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution, and demonstrates its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.

Abstract

Low Earth orbit (LEO) satellite communication, able to provide ubiquitous and continuous connectivity, has become a vital component of future sixth-generation global networks. To further improve service continuity and support spatially non-uniform traffic demand, LEO satellite systems are evolving toward dense multi-shell constellations, where overlapping coverage enables multiple satellites to serve the same traffic region. However, under limited onboard beams, spectrum, and power budgets, such overlap may lead multiple beams to request the same traffic cell, causing duplicate beam-cell requests (DBRs) in downlink scheduling. To address this challenge, we propose a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution. Specifically, we first formulate a long-term scheduling problem that separates pre-resolution beam-cell requests from post-resolution retained transmissions, enabling DBRs to be modeled together with resource block (RB) chunk and power allocation. By leveraging compact local information and per-shell summaries, a typed encoder is then designed to capture heterogeneous service and contention states without relying on a full global map. Furthermore, an autoregressive policy is developed to generate beam-cell, RB chunk, and power-share decisions in accordance with the downlink scheduling sequence, while a deterministic replay-based post-resolution credit mechanism transforms team outcomes into per-beam training signals. Extensive simulation results validate the effectiveness of ShellMean-MAPPO, and demonstrate its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.

Read PDF

Similar papers

Preprint Aug 2026

Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks

The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.

Yassine Afif, Ashutosh Balakrishnan, Philippe Martins et al. · 0 citations
2026

A Heterogeneous Multiagent Reinforcement Learning Approach for Robust Uplink Beamforming in Maritime Satellite Communications

Maritime satellite communications (SATCOMs) are expected to support high-capacity ship-to-satellite uplinks for remote maritime services beyond terrestrial coverage, with low-Earth-orbit (LEO) satellites providing wide-area connectivity. However, robust uplink beamforming in LEO maritime SATCOMs is challenging because...

Huayuan Wang, Bodong Shang, Meixia Tao · 0 citations
2026

Distributed Routing for LEO Satellite Networks: A Multi-Agent Deep Reinforcement Learning Approach With State Information Lag

Multi-agent deep reinforcement learning (MADRL) offers a promising solution for routing in low Earth orbit (LEO) satellite networks. However, large inter-satellite propagation delays lead to severe state information lag in agent interactions, giving rise to decision biases and degraded routing timeliness. To this end,...

Wei-Dan Liu, Tong Liu, Li-Xia Xiao et al. · 0 citations
2026

Collaborative Task Offloading in Space Computing Power Network: A World Model-Based Multi-Agent Reinforcement Learning Approach

Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-...

Yuqi Cong, Zhiwei Wei, Jiarui Chen et al. · 0 citations
2026

Distributed Cooperative Beamforming for Spectrum Sharing in GEO-LEO Heterogeneous Multi-Satellite System

Due to their resilience and global coverage, satellite networks are poised to become a key component for non-terrestrial networks in the future. However, given the scarcity of spectrum resources, the dense deployment of low Earth orbit (LEO) satellites introduces significant interference challenges. Meanwhile, the limi...

Xin Chen, Zhi-Yong Luo · 0 citations
Preprint Aug 2026

LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network Slicing

The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base station mobility, and channel non-stationarity, complicating radio resource management of heterogeneous network slices. In th...

Víctor Vilchez, T. P. C. de Andrade, Edward Hinojosa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.