Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 324-329· 0 citations· 24 references
Abstract
This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.
Unmanned Aerial Vehicles (UAVs) have gained widespread attention in diverse applications like military, medical, aerial surveillance and many more. Presently, the problem of limited bandwidth and geographic factors has raised the need for effective and timely data transfer. Training UAVs with reinforcement learning-based algorithms facilitates autonomous decision-making capabilities. In this paper, we proposed an intelligent system for the optimal UAV selection process by evaluating the continuous performance of each UAV. The analyzing factors are based on the real-world factors affecting the quality of signals, such as noise interference, relative motion between source and wave, and transmission power. Based on the systematic conditions observed, the system provides efficient rewards. To promote the selection of the optimal UAV and enhance the learning process, the state information of the UAV is fed into a deep neural network (DQN), which predicts the 'Q-values'. Our system implements a deep Q-learning algorithm, which enhances the agent's performance by systematically learning from its experience. The model operates accurately by selecting the most reliable UAV, thus, enhancing the throughput by optimal power allocation. It outperforms other conventional models in terms of timely data delivery and energy utilization. The system adapts various complex patterns by analyzing the historical and present scenarios. Empowered by this intelligent system, time-critical decision-making can be achieved with minimal energy consumption.
Divyanshu Bhardwaj, Angel Kanjiya, N. Jadav et al.· 2026 IEEE International Work...· 0 citations
Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.
Aniket Subbanwar, Ojas Joshi, Amit Agarwal· International Conference on...· 0 citations
Unmanned aerial vehicle (UAV) communications are a promising enabler for 6G networks, offering flexible deployment and strong line-of-sight channel conditions. Effective UAV operation requires jointly optimizing trajectory and user scheduling to balance throughput and information freshness. This paper proposes a proximal policy optimization (PPO)-based deep reinforcement learning (DRL) framework that controls UAV movement and user scheduling together via a joint MultiDiscrete action space. We formulate a Markov decision process for a 8-user, $1000 \times 1000 \mathrm{~m}^{2}$ service area with a 3GPP TR 36.777-compliant channel model, where the agent selects both its next position and which user to serve at each time slot. The proposed PPO policy achieves 85.75 Mbps mean throughput, a 24.4% improvement over the AoI-greedy baseline, while reducing mean AoI by 87.7% compared to the throughput-greedy baseline, reaching a Pareto-optimal trade-off between the two competing objectives. An ablation study over the AoI penalty weight confirms a clear throughput-AoI trade-off, validating the joint design.
Quang Tuan Do, Tung Son Do, Thanh Phung Truong et al.· International Conference on...· 0 citations
Most of the existing UAV-assisted communication networks provide service only for static users or deterministically moving ones. In fact, for some complex and dynamically changing scenarios, the users communicating to the UAV may move randomly, with unpredictable mobility. The uncertainty of users’ movements poses a challenge to the guarantee of stable network performance. To tackle this, the paper investigates a UAV-assisted communication network, where a UAV provides communication service for ground users which are moving randomly. We collectively factor in ground user mobility, task duration, and UAV flight restrictions to design precise 3D trajectory for UAV, and formulate them into an optimization problem, aiming to maximize the network throughput while minimizing UAV energy consumption. Considering the dynamics caused by users’ uncertain movement, we transform the optimization problem into a Markov decision process (MDP), then improve the twin-delayed deep deterministic policy gradient (TD3) to design UAV’s 3D trajectory. By utilizing the prior knowledge to accelerate the exploration efficiency, we propose a trajectory design algorithm based on prior knowledge-TD3 (PKTD3-TD), enabling UAV to autonomously adjust flight parameters by leveraging environmental observations under dynamic conditions for enhancing flexibility and intelligence. Simulation results show that our proposed scheme outperforms the compared ones in terms of communication link quality, network throughput and UAV’s energy consumption.
Min Li, M. Dong, Hong Wang et al.· IEEE Transactions on Network...· 0 citations
Reliable link maintenance is currently a critical bottleneck for unmanned aerial vehicle (UAV) swarm communications in complex electromagnetic environments where UAVs encounter both external malicious jamming and internal interference. Most recent studies have treated trajectory design and resource scheduling as decoupled problems or employed standard deep reinforcement learning methods to handle static spectral scenarios. However, these approaches lead to frequent link breakages and slow convergence when dealing with dynamic topologies and spatiotemporal interference. To tackle this challenge, we proposes a joint spatial-spectral adaptive coordination (JSSAC) framework and a deep recurrent attentionbased Q-network (DARQN) approach, utilizing a multi-head attention mechanism to intelligently aggregate heterogeneous neighbor features, thereby enhancing the swarm's adaptability to dynamic network topology. Moreover, considering that the spatial distribution of drones fundamentally determines the upper bound of the signal quality, we designed a communicationaware potential field mechanism that incorporates real-time signal-to-interference-plus-noise ratio feedback. Simulation results demonstrate that compared to DQN and DRQN algorithms, the proposed algorithm achieves transmission success rates of over 92%, representing improvements of 17% and 8% respectively, while also accelerating convergence speed.
Miao Liu, Nan Qi, Hua Jiang et al.· International Mediterranean...· 0 citations