Quantum Enhanced DDPG for Joint Power Spectrum Allocation in D2D-NOMA
Abstract
Power spectrum allocation in Device to Device (D2D) communication using Non-Orthogonal Multiple Access (NOMA) presents a challenging optimization problem due to subchannel pairing, continuous power control, and Successive Interference Cancellation (SIC) ordering. These interdependent parameters result in a mixed-integer, non-convex problem subject to requirement of Quality of Service (QoS) constraints. Existing schemes exhibit limitations, as Deep Q-Networks (DQN) approach restricts from limited action space, leading to suboptimal transmit power allocation and reduced energy efficiency. However, Deep deterministic policy gradient (DDPG) scheme often unstables near SIC threshold. To handle these limitations, this research paper addresses Quantum enhanced DDPG (QDDPG) scheme, which integrates hybrid actor-critic with a feasibility aware projection to enforce SIC and QoS constraints. QDDPG reaches a return of 0.97 in 350 episodes, however DDPG and DQN reach to 0.84 and 0.62, respectively. With 60 D2D pairs, QDDPG attains a sum rate of 9.6 versus 8.7 in DDPG and 7.4 in DQN. Energy efficiency equals 5.8 bits/J at 10 pairs in QDDPG, and 4.7 bits/J and 4.1 bits/J in DDPG and DQN, respectively. These results indicate that the proposed QDDPG shows consistent performance improvements over DDPG and DQN schemes under the considered network conditions.