Skip to content
Conference

Multi-Agent Reinforcement Learning for Dynamic Satellite Constellation Management

Aug 2026 · 2026 International Conference on Modern Sustainable Systems (CMSS) · pp. 125-130 · 0 citations · 20 references

Abstract

Low Earth Orbit (LEO) mega-constellations are becoming central to global broadband delivery, yet their scale and constantly shifting topology make centralized, ground-based control impractical. This paper proposes a multi-agent reinforcement learning (MARL) framework in which each satellite acts as an autonomous agent that jointly balances coverage, latency, collision risk, and energy use. The design follows a centralizedtraining, decentralized-execution (CTDE) paradigm, combining an attention-based inter-satellite communication mechanism with an actor-critic learning scheme, so that each agent conditions its local decisions on neighboring satellite states without relying on a ground controller. A composite reward penalizes fuel waste, collision risk, and communication failure while rewarding coverage and successful relay. The framework was evaluated in a simulated constellation of 40 satellites spanning five orbital regimes and ten ground stations under stochastic debris, weather, and fuel constraints. Over 100 training episodes, the collisionavoidance rate improved from approximately 18% to 85.9%, coverage efficiency stabilized near 62.7%, and average latency fell to 59-62 ms, outperforming Dijkstra routing, Q-Learning, DQN, PPO, MAPPO, and MADDPG while scaling linearly with constellation size. These results position MARL as a practical, scalable foundation for autonomous mega-constellation operations.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.