Skip to content
Conference

Transformer-Enabled Constrained Deep Reinforcement Learning for Joint Satellite Selection and Power Control in LEO-GEO Coexistence Networks

Aug 2026 · 2026 IEEE/CIC International Conference on Communications in China (ICCC) · pp. 683-688 · 0 citations · 14 references

Abstract

In recent years, with the large-scale deployment of low Earth orbit (LEO) satellites, the scale of multi-user uplink transmission in satellite networks has been continuously expanding. During the process of user-satellite association and cochannel spectrum reuse, it not only triggers significant cross-user interference but also poses interference threats to geostationary Earth orbit (GEO) systems. Existing traditional optimization methods suffer from high computational overhead in dynamic topologies, while conventional deep reinforcement learning (DRL) methods fail to effectively model complex link interactions and often encounter sparse rewards during constraint handling. To address these issues, this paper proposes a Pair-Token Transformer-Enabled Constrained Deep Reinforcement Learning (PTC-DRL) method to jointly optimize satellite selection and power control in LEO-GEO coexistence networks under the dual constraints of GEO interference avoidance and minimum elevation angles, aiming for system sum-rate maximization. By taking the usersatellite association as the basic interaction unit, the proposed method leverages a Transformer to capture cross-user and cross-satellite global spatial coupling, effectively enhancing state representation capabilities in complex interference scenarios. Furthermore, a feasibility-aware dual-bucket prioritized experience replay mechanism is designed to separately manage feasible and infeasible samples through adaptive sampling, thereby improving sample utilization and constraint learning efficiency. Simulation results show that PTC-DRL achieves stable dualconstraint satisfaction across varying user scales while delivering superior system throughput in medium-to high-load scenarios. Particularly in the 20-user case, the proposed algorithm improves the system sum-rate by $\mathbf{2. 8 7 \%}$ and $\mathbf{1 0. 0 5 \%}$ compared with the best-performing baseline schemes, Lagrangian Double DQN and drIBPA, respectively, demonstrating its effectiveness in high-load resource allocation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.