A Differentiable Optimization Enhanced Actor-Critic Framework for Resource Allocation in Spectrum Sharing UAV Networks
Abstract
Spectrum sharing is promising to alleviate the spectrum scarcity problem of the unmanned aerial vehicle (UAV) networks. Resource allocation is also of crucial importance to improve the spectrum efficiency and network performance. However, the resource allocation problems are typically NPhard when the number of the optimized resource variables and that of the users are large, making traditional optimizationbased approaches computationally unaffordable for real-time scenarios. As an alternative, deep reinforcement learning (DRL)based approaches have gained significant attention due to their flexibility and end-to-end capability in solving such complex resource allocation problems. However, these approaches typically rely on penalty-based reward shaping to handle constraints, leading to policies that struggle to satisfy hard constraints and exhibit limited generalization and robustness to environmental changes. To address these issues, a differentiable optimization enhanced actor-critic framework is proposed. The key is to embed a differentiable optimization layer within the actor network, integrating prior knowledge with the exploration and end-toend learning capabilities of DRL. By exploiting the proposed framework, an intelligent spectrum allocation and power control scheme is developed. Simulation results demonstrate that our proposed scheme is superior to the benchmark schemes in terms of convergence performance and exhibits better generalization and robustness to external disturbances and system parameter changes.