Skip to content
Conference

Reinforcement Learning-Based Decode-and-Forward UAV Relay Trajectory Optimization

Jul 2026 · International Conference on Signal Processing and Communications · pp. 1-5 · 0 citations · 6 references

Abstract

Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.

View source

Similar papers

Open access 2026

Joint Trajectory and Power Optimization for UAV-Relay: A Constraint Handling Approach with the Bat Algorithm

— Unmanned Aerial Vehicles (UAVs) have emerged as flexible relay platforms capable of enhancing wireless connectivity in beyond-5G and 6G networks. This paper investigates the joint optimization of UAV trajectory and power allocation to maximize end-to-end throughput under practical mobility and power constraints. The problem is highly non-convex due to the strong coupling between trajectory variables and transmission power. To address this challenge, we develop a penalty-based metaheuristic framework that incorporates a constraint-handling mechanism into the Bat Algorithm (BAT). Simulation results show that the proposed BAT-based approach achieves significant throughput improvement, efficient power allocation, and fast convergence compared with baseline convex optimization and heuristic schemes. These findings highlight the potential of BAT for reliable and energy-efficient UAV-assisted communication in future wireless networks.

Pham Thi Quynh Trang · 0 citations
Conference Jul 2026

Joint Trajectory and Scheduling Optimization for UAV-Assisted 6G Networks: A Deep Reinforcement Learning Approach with Throughput–AoI Trade-off

Unmanned aerial vehicle (UAV) communications are a promising enabler for 6G networks, offering flexible deployment and strong line-of-sight channel conditions. Effective UAV operation requires jointly optimizing trajectory and user scheduling to balance throughput and information freshness. This paper proposes a proximal policy optimization (PPO)-based deep reinforcement learning (DRL) framework that controls UAV movement and user scheduling together via a joint MultiDiscrete action space. We formulate a Markov decision process for a 8-user, $1000 \times 1000 \mathrm{~m}^{2}$ service area with a 3GPP TR 36.777-compliant channel model, where the agent selects both its next position and which user to serve at each time slot. The proposed PPO policy achieves 85.75 Mbps mean throughput, a 24.4% improvement over the AoI-greedy baseline, while reducing mean AoI by 87.7% compared to the throughput-greedy baseline, reaching a Pareto-optimal trade-off between the two competing objectives. An ablation study over the AoI penalty weight confirms a clear throughput-AoI trade-off, validating the joint design.

Quang Tuan Do, Tung Son Do, Thanh Phung Truong et al. · 0 citations
2026

Hybrid DRL-Based Sensing and Age of Information Optimization for UAV-Enabled ISCC

In low-altitude economy (LAE), the deployment of unmanned aerial vehicles (UAVs) provides substantial convenience and enhances operational efficiency. This letter investigates a joint resource and trajectory optimization problem in a UAV-enabled integrated sensing, computation, and communication (ISCC) network where the UAV senses and processes the status information from sensing targets (STs), and then sends the computed results to the data collection center (DC). Aiming to maximize the sensing data volume while minimizing the age of information (AoI), we jointly optimize the UAV’s sensing schedule, number of sensing trials, time allocation, transmit power, CPU frequency, and trajectory. The problem is formulated as a Markov decision process (MDP), and a hybrid deep reinforcement learning (DRL) framework is proposed to derive optimal policies. Specifically, we adopt the twin delayed deep deterministic policy gradient (TD3) framework and enhance it with a hybrid-baseline prioritized experience replay (PER) mechanism, denoted as HPTD3. Simulation results demonstrate that the proposed approach significantly outperforms benchmark schemes.

Honghao Qi, Muqing Wu, Zijian Zhang et al. · 0 citations
Preprint Jul 2026

CRB-Driven Beamforming and Trajectory Optimization for UAV-assisted ISAC System

In this paper, we study an unmanned aerial vehicle (UAV)-assisted integrated sensing and communication (ISAC) system, where a UAV enhances the sensing capability of a base station (BS) towards a target while ensuring reliable communication towards a downlink user. This architecture is practically attractive for future wireless networks due to the UAV's controllable mobility and adaptive sensing coverage in wireless environments. The sensing performance is characterized by the average Cram\'er-Rao bound (CRB), which quantifies the minimum variance of the unbiased angle-of-arrival estimation. To enhance the sensing performance, the UAV trajectory and beamforming parameters are jointly optimized under power and mobility constraints, while satisfying communication requirements to the downlink user. To address the resulting non-convex problem, we employ null-space projection for beamforming design and adopt deep reinforcement learning for the trajectory optimization over a discrete-time scale. In each time slot, beamforming is optimized based on the channel state information to improve CRB performance while mitigating interference between the BS and the communication user. Simulation results demonstrate that the proposed method significantly reduces the time-averaged CRB by over 10%, compared with the ISAC system without UAV assistance, and also achieves a higher sensing accuracy than both the fixed-UAV-trajectory and the maximum-ratio-transmission-based beamforming benchmarks.

Yi Yang, Qianqian Zhang, Huaxia Wang · 0 citations
2026

3-D Trajectory Design Based on Deep Reinforcement Learning for UAV-Assisted Communication Networks

Most of the existing UAV-assisted communication networks provide service only for static users or deterministically moving ones. In fact, for some complex and dynamically changing scenarios, the users communicating to the UAV may move randomly, with unpredictable mobility. The uncertainty of users’ movements poses a challenge to the guarantee of stable network performance. To tackle this, the paper investigates a UAV-assisted communication network, where a UAV provides communication service for ground users which are moving randomly. We collectively factor in ground user mobility, task duration, and UAV flight restrictions to design precise 3D trajectory for UAV, and formulate them into an optimization problem, aiming to maximize the network throughput while minimizing UAV energy consumption. Considering the dynamics caused by users’ uncertain movement, we transform the optimization problem into a Markov decision process (MDP), then improve the twin-delayed deep deterministic policy gradient (TD3) to design UAV’s 3D trajectory. By utilizing the prior knowledge to accelerate the exploration efficiency, we propose a trajectory design algorithm based on prior knowledge-TD3 (PKTD3-TD), enabling UAV to autonomously adjust flight parameters by leveraging environmental observations under dynamic conditions for enhancing flexibility and intelligence. Simulation results show that our proposed scheme outperforms the compared ones in terms of communication link quality, network throughput and UAV’s energy consumption.

Min Li, M. Dong, Hong Wang et al. · 0 citations
Conference Jul 2026

Joint Trajectory and Spectrum Optimization for Anti-Jamming UAV Swarms: A DRL Approach

Reliable link maintenance is currently a critical bottleneck for unmanned aerial vehicle (UAV) swarm communications in complex electromagnetic environments where UAVs encounter both external malicious jamming and internal interference. Most recent studies have treated trajectory design and resource scheduling as decoupled problems or employed standard deep reinforcement learning methods to handle static spectral scenarios. However, these approaches lead to frequent link breakages and slow convergence when dealing with dynamic topologies and spatiotemporal interference. To tackle this challenge, we proposes a joint spatial-spectral adaptive coordination (JSSAC) framework and a deep recurrent attentionbased Q-network (DARQN) approach, utilizing a multi-head attention mechanism to intelligently aggregate heterogeneous neighbor features, thereby enhancing the swarm's adaptability to dynamic network topology. Moreover, considering that the spatial distribution of drones fundamentally determines the upper bound of the signal quality, we designed a communicationaware potential field mechanism that incorporates real-time signal-to-interference-plus-noise ratio feedback. Simulation results demonstrate that compared to DQN and DRQN algorithms, the proposed algorithm achieves transmission success rates of over 92%, representing improvements of 17% and 8% respectively, while also accelerating convergence speed.

Miao Liu, Nan Qi, Hua Jiang et al. · 0 citations