DEEP REINFORCEMENT LEARNING-BASED DYNAMIC SPECTRUM ACCESS FOR 6G HETEROGENEOUS COGNITIVE RADIO NETWORKS
Abstract
The proliferation of heterogeneous radio access technologies in sixth-generation (6G) wireless networks demands a fundamental rethinking of spectrum management strategies. Traditional spectrum sensing approaches, designed for relatively static channel conditions, are inadequate for the dynamic, interference-rich environments that characterize 6G deployments spanning sub-6 GHz, millimetre-wave, and terahertz bands simultaneously. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users (SUs) learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models. Specifically, a Double Deep Q-Network (DDQN) architecture is adopted, augmented with a prioritized experience replay mechanism that accelerates policy convergence under non-stationary channel conditions. The proposed agent observes a composite state space encoding instantaneous channel occupancy, signal-to-interference-plus-noise ratio (SINR), primary user (PU) activity patterns, and residual energy levels, and selects actions that jointly optimize spectrum utilization efficiency, interference avoidance, and energy consumption. Simulation experiments conducted over a heterogeneous network topology with four primary users and eight secondary users demonstrate that the proposed DDQN-based scheme achieves a throughput gain of approximately 34% over conventional energy detection-based sensing, reduces interference to primary users by 61%, and attains a detection probability of 0.94 at a false alarm rate of 0.05. These results confirm the practical viability of DRL as a spectrum management backbone for next-generation cognitive radio systems.