Skip to content
Open access

MEMORY-AUGMENTED REINFORCEMENT LEARNING FOR UAV NAVIGATION USING PPO-LSTM

Aug 2026 · Kufa journal of Engineering · 0 citations · 5 references

TL;DR

Experimental outcomes show that the PPO-LSTM described herein achieves smoother paths, more robust reward convergence, and a much lower rate of collision than regular PPO, and generalizes to new environments with movable obstacles.

Abstract

The problems of partial observability and sensor shortage pose a significant challenge for autonomous Unmanned Aerial Vehicles (UAVs) as they prove to be challenging for conventional Deep Reinforcement Learning (DRL) methods to undertake well under such conditions. In this paper, a memory-augmented Proximal Policy Optimization (PPO) model extended using a Long Short-Term Memory (LSTM) network is proposed as a solution to such challenges. The observation space is constructed from 2D LiDAR and Inertial Measurement Unit (IMU) data to sense simultaneously external observation and internal state of motion, whereas the action space consists of continuous velocity commands. A shaped reward function is optimized for encouraging safe target approaching, obstacle avoidance, and convergence speed. Experimental outcomes show that the PPO-LSTM described herein achieves smoother paths, more robust reward convergence, and a much lower rate of collision than regular PPO. It also generalizes to new environments with movable obstacles. Qualitatively, the success rate increased from 64.5% to 83.9%, collision frequency reduced by over 70%, and path efficiency increased from 0.60 to 0.85, without suffering from unstable training behavior

Read PDF

Similar papers

Preprint Aug 2026

PILOT: Privileged Imitation Learning for End-to-End Motion Planning of Autonomous UAVs under Partial Observability

PILOT, a constraint-aware privileged imitation learning framework for vision-based end-to-end UAV motion planning under partial observability, is presented, demonstrating the practical feasibility and cross-domain generalization of the planner.

Qing-Rui Zhang, Feng Xue, Xiang Zhou et al. · 0 citations
Conference Aug 2026

Teacher-Guided Asymmetric Reinforcement Learning for End-to-End Visual Navigation of UAVs

Autonomous navigation of low-altitude unmanned aerial vehicles (UAVs) in cluttered environments is challenging due to partial observability, limited onboard perception, and inefficient exploration in end-to-end reinforcement learning. This paper proposes a teacher-guided asymmetric reinforcement learning framework for...

Yi-Min Wei, Qiu-Quan Guo, Cai-Zheng Wang et al. · 0 citations
Aug 2026

Deep reinforcement learning–based safe path planning for leader–follower robots

This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.

Ehsan Kazemi Tameh, Mohammadreza Estarki, Saeed Khodaygan · 0 citations
Preprint Aug 2026

Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboar...

M. E. Deowan, Eleni Kelasidi · 0 citations
Open access Aug 2026

Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation

This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshe...

Yi-Tan Wang, Yang-Yang Zhang, Zhen-Xing Gao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.