An improved DRL algorithm, Dropout-based Prioritized Soft Actor-Critic (DPSAC), which integrates the Soft Actor-Critic (SAC) algorithm with Prioritized Experience Replay (PER) and the Dropout technique is proposed, and two innovative approaches are introduced to enhance the algorithm’s performance.
Abstract
In uncrewed aerial vehicle (UAV)-assisted Internet of Things (IoT) networks, UAVs often need to fly at low altitudes and navigate through obstacles to maintain reliable communication with IoT nodes during data collection missions. This paper proposes a novel deep reinforcement learning (DRL)-based approach for 3D UAV path planning and obstacle avoidance, with the objective of minimizing data collection time from IoT nodes distributed across complex urban environments. To address this problem, we propose an improved DRL algorithm, Dropout-based Prioritized Soft Actor-Critic (DPSAC), which integrates the Soft Actor-Critic (SAC) algorithm with Prioritized Experience Replay (PER) and the Dropout technique. Furthermore, two innovative approaches are introduced to enhance the algorithm’s performance. First, the Episodic Environment (EN) training approach introduces random variations in obstacle number, position, and height across training episodes, thereby enhancing the agent’s ability to generalize its learned policy to new and unknown environments. Second, the Switching Reward mechanism reduces penalties for collisions and boundary violations in the reward function during the early stages of training, thereby facilitating exploration and accelerating the agent’s learning of IoT-related tasks. Simulation results demonstrate that the proposed DRL-based approach achieves faster convergence and greater stability during the training process compared to baseline algorithms. Specifically, experiments conducted in new and complex environments show that this method can collect data with an average collision-free success rate of 98% from 10 IoT nodes and 95% from 20 IoT nodes, confirming its remarkable superiority over the baseline algorithms.
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory pla...
Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al.· arXiv.org· 0 citations
Experimental results validate the effectiveness of the proposed reinforcement learning approach based on the Deep Q-Network algorithm, demonstrating its superior obstacle avoidance capabilities and exceptional energy optimization performance in complex urban settings.
Yang Li, Xinjie Qian, Yanxiu Wang et al.· International Conference on...· 0 citations
To address the challenge of rapid and precise obstacle avoidance for unmanned aerial vehicles (UAVs) in complex urban environments, rugged canyons, and other unstructured environments, this paper proposes a vision-based navigation algorithm. By combining the strengths of deep reinforcement learning (DRL) and convolutio...
Dongliang Wang, Yong-Qiang Jin, Weicheng Luo et al.· Italian National Conference...· 0 citations
Urban fire rescue poses severe challenges to the real-time performance and obstacle avoidance capabilities of unmanned aerial vehicle (UAV) path planning. Existing methods (such as A*, RRT, and standard DQN) have problems such as low search efficiency, insufficient obstacle avoidance ability, or slow convergence in com...
Rui Qin, Han-Jing Zhou· International Conference on...· 0 citations
In recent years, Unmanned Aerial Vehicles (UAVs) have gradually been widely used in various fields such as regional search and disaster relief, and the development of related technologies has also experienced unprecedented growth. Compared to individual UAVs, the collaborative execution of tasks by UAV swarms has more...
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.