D3QN-Based Joint Power and Reflection Optimization for Energy-Efficient Anti-Jamming against Dynamic Jamming Attacks in Backscatter-Enabled 6G IoT Networks
Abstract
This paper develops a Dueling Double Deep Q-Network (D3QN) framework for joint transmit-power and reflection-coefficient control in backscatter-enabled 6G Internet-of-Things networks subject to random, reactive, and time-varying jamming. The problem is formulated as a jammer-aware Markov decision process with a 20-action power–reflection space and a normalized reward that balances spectral efficiency, transmit-power consumption, and low-SINR outage. The proposed method is evaluated against DQN, DDQN, PPO, TD3, SAC, random selection, and fixed control over five independent seeds with 95% confidence intervals. Under random jamming at Pj = 30 dBm, D3QN improves spectral efficiency by 19.83% and reduces outage probability by 6.99% relative to DQN. Under a common dynamic-jamming trace, it provides a 2.12% spectral-efficiency gain while using an average transmit power of 0.427 W, with the maximum-power action selected in only 17.6% of the slots. The 20-action design also improves spectral efficiency by 4.21% over a 9-action grid while using 59.18% fewer outputs than a 49-action grid. Under strong reactive jamming, D3QN improves energy efficiency over DQN by 2.54%, but exhibits lower spectral efficiency and slightly higher outage, revealing a rate - energy - reliability trade -off. Overall, the proposed framework offers a practical balance among communication performance, energy use, and action-space complexity