This paper develops a Dueling Double Deep Q-Network (D3QN) framework for joint transmit-power and reflection-coefficient control in backscatter-enabled 6G Internet-of-Things networks subject to random, reactive, and time-varying jamming. The problem is formulated as a jammer-aware Markov decision process with a 20-action power–reflection space and a normalized reward that balances spectral efficiency, transmit-power consumption, and low-SINR outage. The proposed method is evaluated against DQN, DDQN, PPO, TD3, SAC, random selection, and fixed control over five independent seeds with 95% confidence intervals. Under random jamming at Pj = 30 dBm, D3QN improves spectral efficiency by 19.83% and reduces outage probability by 6.99% relative to DQN. Under a common dynamic-jamming trace, it provides a 2.12% spectral-efficiency gain while using an average transmit power of 0.427 W, with the maximum-power action selected in only 17.6% of the slots. The 20-action design also improves spectral efficiency by 4.21% over a 9-action grid while using 59.18% fewer outputs than a 49-action grid. Under strong reactive jamming, D3QN improves energy efficiency over DQN by 2.54%, but exhibits lower spectral efficiency and slightly higher outage, revealing a rate - energy - reliability trade -off. Overall, the proposed framework offers a practical balance among communication performance, energy use, and action-space complexity
Lê Hoàng Hiệp, Huy Huu Ngo· Journal of Science and Techn...· 0 citations
—This paper proposes an adaptive multi-mode Deep Reinforcement Learning (DRL) framework for intelligent RIS-assisted anti-jamming communication in dynamic 6G wireless networks. The proposed Framework jointly integrates RIS beamforming, channel hopping, and transmit power adaptation through a DRL-Driven decision engine capable of dynamically responding to varying interference conditions and channel fluctuations. To improve deployment realism, practical constraints including imperfect Channel State Information (CSI), finite-resolution RIS phase quantization, reflection loss, control delay, and user mobility are incorporated into the system model. The anti-jamming problem is formulated as a Markov decision process and solved using DQN, PPO, and SAC algorithms. Extensive simulations are conducted using MATLAB-based wireless channel modeling and Python-based DRL training platforms. Simulation results demonstrate that the proposed framework achieves approximately 25%–40% higher throughput and 18%– 35% SINR improvement compared with conventional anti-jamming approaches. Moreover, the proposed scheme maintains stable communication performance under strong jamming power, CSI uncertainty, and high-mobility scenarios. Statistical evaluations over 20 independent random seeds further confirm the robustness and reproducibility of the proposed framework.
Lê Hoàng Hiệp, Huu-Huy Ngo· Journal of Communications So...· 1 citation