A physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques, and closes the loop through a high-fidelity Simulink environment is investigated.
Abstract
Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only four available control inputs must govern six degrees of freedom. This paper investigates a physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques $(T, \tau_x, \tau_y, \tau_z)$, and closes the loop through a high-fidelity Simulink environment. Our simulator integrates a 12-state rigid-body model (MATLAB Level-2 S-Function) with (i) an Action2RPM allocation based on the Moore-Penrose pseudo-inverse of a coefficient matrix derived from thrust and drag terms, and (ii) first-order actuator dynamics for each motor (time constant $T_m = 0.076$ s), including rotor gyroscopic coupling. A shaped reward balances goal-reaching and stability using an exponential position well, attitude penalties, and quadratic velocity costs. Four DRL algorithms, DDPG, TD3, PPO, and SAC, are evaluated in two stages: (S1) thrust-only hover and (S2) hover with pitch torque and a translated goal. Results show that SAC and TD3 achieve superior stability and exploration efficiency, while PPO is less sample-efficient. The study highlights the significance of modeling actuator lags and aerodynamic moments for stable low-level control and provides a reproducible benchmark for quadcopter DRL.
This paper explores and evaluates both traditional and learning-based methods for spacecraft control and coupled robotic arm manipulation in a microgravity environment. Space junk, debris, and out-of-control satellites currently floating in orbit pose a severe risk to critical space assets, necessitating debris removal...
Patrick Coulon, Charles East, Luke Busse et al.· National Aerospace and Elect...· 0 citations
These findings support the feasibility of native Rust-based D3QN training for real-time fixed-wing simulation, while the observed inter-seed variability indicates that reward shaping and convergence robustness require further evaluation.
Saugat Chaudhary Tharu, Shrutika Ojha, Rija Bhomi et al.· Journal of Advances in Mathe...· 0 citations
This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupl...
Zied Ben Hazem, Muhammed Özdemir, Firas Saidi et al.· Discover Robotics· 1 citation
Yaw regulation of biomimetic underwater robots is complicated by flexible body motion, nonlinear hydrodynamics, and coupled actuation. This study examines whether a policy trained in simulation can be deployed on an existing robotic sea lion (RSL) without changing its hardware or low-level controllers. A deep determini...
Ze-Yi Zhang, Yu-Hong Liu, Shuang-Tao Liu et al.· Journal of Marine Science an...· 0 citations
This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot...
Filip Połatyński, Paweł Skruch· International Conference on...· 0 citations
A reinforcement-learning-based training framework for the hierarchical control architecture of TSFV-UAVs is proposed, which introduces an adaptive guidance mechanism and a varying-gradient reward function to improve training convergence.