Most of the existing UAV-assisted communication networks provide service only for static users or deterministically moving ones. In fact, for some complex and dynamically changing scenarios, the users communicating to the UAV may move randomly, with unpredictable mobility. The uncertainty of users’ movements poses a challenge to the guarantee of stable network performance. To tackle this, the paper investigates a UAV-assisted communication network, where a UAV provides communication service for ground users which are moving randomly. We collectively factor in ground user mobility, task duration, and UAV flight restrictions to design precise 3D trajectory for UAV, and formulate them into an optimization problem, aiming to maximize the network throughput while minimizing UAV energy consumption. Considering the dynamics caused by users’ uncertain movement, we transform the optimization problem into a Markov decision process (MDP), then improve the twin-delayed deep deterministic policy gradient (TD3) to design UAV’s 3D trajectory. By utilizing the prior knowledge to accelerate the exploration efficiency, we propose a trajectory design algorithm based on prior knowledge-TD3 (PKTD3-TD), enabling UAV to autonomously adjust flight parameters by leveraging environmental observations under dynamic conditions for enhancing flexibility and intelligence. Simulation results show that our proposed scheme outperforms the compared ones in terms of communication link quality, network throughput and UAV’s energy consumption.
Min Li, M. Dong, Hong Wang et al.· IEEE Transactions on Network...· 0 citations
The coexistence of Ultra-Reliable Low-Latency Communication (URLLC) and Enhanced Mobile Broadband (eMBB) can simultaneously meet the reliability and real-time performance requirements of critical services as well as the requirements of high-bandwidth services. In 5G and beyond 5G (B5G) networks, radio access network (RAN) slicing can ensure the differentiated quality of service (QoS) for coexisting heterogeneous services. This paper investigates a dynamic resource allocation framework, aiming to maximize resource utilization subject to QoS constraints. The optimization problem is an NP-hard integer program. To address this issue, we propose a hierarchical proximal policy optimization (PPO) model assisted by a traffic prediction algorithm, namely graph-aggregated Extended Long Short-Term Memory (GxLSTM), which decouples the resource allocation behavior between eMBB and URLLC. Transfer learning is introduced to improve the learning efficiency, forming a Transfer Learning-assisted Hierarchical PPO algorithm (TL-HPPO). In addition, considering the lightweight algorithm deployment requirements for edge networks, based on angular knowledge distillation (AKD), GxLSTM and TL-HPPO are respectively distilled to obtain AKD-traffic prediction (AKD-TP) and AKD-resource allocation (AKD-RA), thereby reducing computational complexity. Simulation results show that the proposed algorithms outperform comparative algorithms while ensuring QoS, effectively improving resource utilization.
Yixuan Bai, Heng Wang· IEEE Transactions on Communi...· 0 citations