A Multimodal Deep Reinforcement Learning Framework for Autonomous UAV Navigation in Gazebo–ROS Environments
Abstract
Unmanned Aerial Vehicle (UAV) autonomous navigation is a key capability for UAVs, particularly when in complex and dynamic environments where continuous human control is impractical. While commonly used rule-based navigation and path-planning techniques can be effective in structured environments, they can be ineffective when the environment is unknown, obstacles are moving, or sensor data is noisy. A UAV navigation framework based on DRL methods, including DQN, PPO, and the Actor-Critic algorithm, is presented. The proposed system is developed and evaluated in a Gazebo-ROS simulator with multimodal sensor input: camera, LiDAR, and IMU. UAV navigation is modeled as a Markov decision process (MDP), and the agent learns navigation decisions through continuous interaction with the environment. The success rate, collision rate, cumulative reward, path length, time to goal, and convergence episodes are used to compare the performance of the three DRL algorithms. Results show that PPO outperforms the other algorithms in every aspect, having the highest success rate, the lowest number of collisions, the highest convergence rate, and the smoothest navigation behavior. Actor–Critic can also be well adapted to a dynamic environment, whereas DQN is less suitable for continuous UAV control due to its discrete action space. The results show that DRL-based navigation is feasible in simulation and reveal the research challenges UAVs will face in the real world.