Quantum Deep Reinforcement Learning On-the-Fly: An Energy Efficient Scheme for Autonomous Aerial Vehicles
Abstract
Intelligent Reflecting Surface (IRS)-assisted Autonomous Aerial Vehicle (AAV) networks have emerged as a promising paradigm for enhancing coverage and spectral efficiency in next-generation wireless systems. However, the high mobility in AAV, severe channel variations, and stringent energy constraints make joint trajectory and IRS optimization a challenging problem. To mitigate this issue, in this paper, we propose an energy-efficient IRS-assisted AAV communication framework which jointly optimizes AAV trajectory and IRS phase-shift configuration to maximize communication performance under realistic rotary-wing propulsion energy consumption. The problem is formulated as a Markov Decision Process (MDP) capturing the coupling between AAV mobility, wireless channels, and IRS control. To address the resulting non-convex optimization, a Quantum Deep Deterministic Policy Gradient (QDDPG) algorithm is designed, by integrating variational quantum circuits with actor-critic reinforcement learning for continuous action control. The proposed framework improves exploration capability and policy representation in high-dimensional state-action spaces. Simulation results represent the superiority of QDDPG with respect to energy efficiency (EE), fast convergence and throughput as compared to classical DRL baselines such as DDPG and DDQN schemes.