Skip to content
Open access

Physics-constrained multimodal reinforcement learning for local UAV Navigation in complex static obstacle environments

Jul 2026 · Scientific Reports · Vol 16 · 0 citations · 35 references
Medicine

TL;DR

PC-MM-SAC is presented, a task-oriented multimodal reinforcement-learning framework with a lightweight physics-constrained action-execution layer for local navigation in complex static obstacle environments that achieves the highest success rate among the evaluated methods under the current task setting.

Abstract

Local navigation in complex static obstacle environments requires an unmanned aerial vehicle to reach a target while avoiding obstacles under limited local perception. U-shaped non-convex obstacles acting as local traps further increase the difficulty of detour selection and continuous-control execution. A key challenge is to jointly capture forward scene structure, surrounding geometric constraints, and execution-level control requirements within a unified decision-making pipeline. To address this challenge, we present PC-MM-SAC, a task-oriented multimodal reinforcement-learning framework with a lightweight physics-constrained action-execution layer for local navigation in complex static obstacle environments. The method learns from structured observations built from depth maps, LiDAR measurements, and task-related state variables. The layer applies bounded action mapping and yaw-rate variation limiting to improve the smoothness and continuity of continuous-control commands without introducing an additional online optimization-based controller. We further employ an event-conditioned modality value-sensitivity analysis during evaluation to characterize how the critic’s local sensitivity to different information sources varies across representative navigation phases. In AirSim experiments, PC-MM-SAC achieves a success rate of 0.88 and an episode reward of 772, attaining the highest success rate among the evaluated methods under the current task setting. Ablation and behavioral analyses suggest that multimodal observations are the primary contributor to the observed performance improvement, while the physics-constrained action-execution layer is associated with smoother trajectory execution and reduced local oscillation.

Read PDF

Similar papers

Preprint Sep 2026

Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors

We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) operating at a fixed altitude in unknown dynamic environments, where real-time avoidance of both static and dynamic obstacles must be achieved under conditions of extreme partial observability. The UAV is con...

Mohammad Reza Ranjbar Divkoti, A. Aguiar · 0 citations
Open access Jul 2026

A Belief-Driven Hybrid Reinforcement Learning Framework for Decentralized Multi-Robot Navigation Under Partial Observability

Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a c...

V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al. · 0 citations
Preprint Aug 2026

Hierarchical Topology-Aware Planning and Control of Underwater Vehicle-Manipulator Systems in Confined Environments

This paper addresses autonomous intervention with an underwater vehicle--manipulator system (UVMS) in confined, cluttered, and partially known environments, where poor maneuverability, narrow passages, and uncertain execution may cause the robot to enter unrecoverable regions. We propose MANTA, a three-layer hierarchic...

Mohamed Abdelwahab, Ruggero Carli, Damiano Varagnolo et al. · 0 citations
Conference Jul 2026

Empirical Evaluation of Proximal Policy Optimization for UAV Navigation in Constrained Indoor Environments

Autonomous navigation of unmanned aerial vehicles (UAVs) in constrained indoor environments remains a challenging problem due to limited maneuvering space and high collision risk. This paper presents an empirical evaluation of a reinforcement learning-based approach for UAV path planning using Proximal Policy Optimizat...

Ashraf Suyyagh, Tasneem Al-Qat, Hala Mukheimer et al. · 0 citations
Aug 2024

LSTP-Nav: Lightweight Spatiotemporal Policy for Map-Free Multi-Agent Navigation With LiDAR

This paper proposes LSTP-Nav, a lightweight, decentralized navigation framework built on LSTP-Net that maps stacked 2D LiDAR observations, goal information, and velocity feedback directly to action and introduces an HS reward to provide smooth, heading-aware safety feedback, and develops PhysReplay-SimLab to improve tr...

Xingrong Diao, Zhi-Qiang Sun, Jian-Wei Peng et al. · 0 citations
Open access Sep 2026

ACR-Nav: Localization-Free Corridor Navigation via Action-Conditioned Scalar-Range Evolution

Mapless navigation often removes global maps while retaining localization-derived goal vectors or bearings. We study a stricter setting in which a mobile robot observes only local LiDAR, scalar goal range, and short histories of executed actions; neither pose nor goal direction is provided to the policy. We introduce A...

Qiguang Shen, Zhao-Yue Wang, Yi-Fei Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.