Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 813-818· 0 citations· 11 references
Abstract
Trajectory planning for unmanned aerial vehicles (UAVs) in dynamic and partially observable environments becomes more complex when extended from two-dimensional to three-dimensional navigation. Although Deep Reinforcement Learning (DRL) methods have shown strong performance in 2D scenarios, their application to 3D spaces requires redesigned observation models, action representations, and safety mechanisms. This paper extends a 2D DRL-based trajectory planning framework to 3D environments using Proximal Policy Optimization (PPO), Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG). UAV agents are trained to reach randomly placed 3D targets while avoiding static and dynamic obstacles using only local sensory information. The observation space combines a local 3D occupancy representation with a relative 3D goal vector, preserving partial observability and avoiding reliance on a global map. This article proposes that the simulation results demonstrate robust, collision-aware navigation and improved safety and trajectory efficiency in each one of the DRL algorithms implemented, each having positive and negative specifics.
This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.
This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboar...
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
Efficient 3-D path planning for autonomous underwater vehicles (AUVs) in dynamic submarine environments presents a significant challenge due to complex seabed terrain, ocean currents, and obstacles. In view of the adaptability and generalization limitations of traditional methods, this article proposes the reward-adapt...
Xin Cheng, Hai Jin, Yun Chen et al.· IEEE Systems Journal· 0 citations
Autonomous navigation of unmanned aerial vehicles (UAVs) in constrained indoor environments remains a challenging problem due to limited maneuvering space and high collision risk. This paper presents an empirical evaluation of a reinforcement learning-based approach for UAV path planning using Proximal Policy Optimizat...
Ashraf Suyyagh, Tasneem Al-Qat, Hala Mukheimer et al.· IEEE Jordan Conference on Ap...· 0 citations
Experimental outcomes show that the PPO-LSTM described herein achieves smoother paths, more robust reward convergence, and a much lower rate of collision than regular PPO, and generalizes to new environments with movable obstacles.
M. Haddad, Dhayaa Khudher· Kufa journal of Engineering· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.