Skip to content

Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator Dynamics

Jul 2026 · arXiv.org · Vol abs/2607.25985 · 0 citations · 26 references
Computer Science Engineering

TL;DR

A physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques, and closes the loop through a high-fidelity Simulink environment is investigated.

Abstract

Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only four available control inputs must govern six degrees of freedom. This paper investigates a physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques $(T, \tau_x, \tau_y, \tau_z)$, and closes the loop through a high-fidelity Simulink environment. Our simulator integrates a 12-state rigid-body model (MATLAB Level-2 S-Function) with (i) an Action2RPM allocation based on the Moore-Penrose pseudo-inverse of a coefficient matrix derived from thrust and drag terms, and (ii) first-order actuator dynamics for each motor (time constant $T_m = 0.076$ s), including rotor gyroscopic coupling. A shaped reward balances goal-reaching and stability using an exponential position well, attitude penalties, and quadratic velocity costs. Four DRL algorithms, DDPG, TD3, PPO, and SAC, are evaluated in two stages: (S1) thrust-only hover and (S2) hover with pitch torque and a translated goal. Results show that SAC and TD3 achieve superior stability and exploration efficiency, while PPO is less sample-efficient. The study highlights the significance of modeling actuator lags and aerodynamic moments for stable low-level control and provides a reproducible benchmark for quadcopter DRL.

View source

Similar papers

Conference Aug 2026

A Comparative Study of MPC and Reinforcement Learning Control for Space Robotics Manipulators

This paper explores and evaluates both traditional and learning-based methods for spacecraft control and coupled robotic arm manipulation in a microgravity environment. Space junk, debris, and out-of-control satellites currently floating in orbit pose a severe risk to critical space assets, necessitating debris removal...

Patrick Coulon, Charles East, Luke Busse et al. · 0 citations
Open access Aug 2026

A Rust-Implemented Dueling Double DQN for Fixed-Wing Autonomous Flight

These findings support the feasibility of native Rust-based D3QN training for real-time fixed-wing simulation, while the observed inter-seed variability indicates that reward shaping and convergence robustness require further evaluation.

Saugat Chaudhary Tharu, Shrutika Ojha, Rija Bhomi et al. · 0 citations
Open access Jul 2026

Deep reinforcement learning for accurate trajectory tracking in 2-DOF helicopter dynamics using twin delayed DDPG

This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupl...

Zied Ben Hazem, Muhammed Özdemir, Firas Saidi et al. · 1 citation
Open access Sep 2026

Sim-to-Real Yaw Control of a Robotic Sea Lion Using Deep Reinforcement Learning

Yaw regulation of biomimetic underwater robots is complicated by flexible body motion, nonlinear hydrodynamics, and coupled actuation. This study examines whether a policy trained in simulation can be deployed on an existing robotic sea lion (RSL) without changing its hardware or low-level controllers. A deep determini...

Ze-Yi Zhang, Yu-Hong Liu, Shuang-Tao Liu et al. · 0 citations
Conference Aug 2026

End-to-End Control of a Quadruped Robot Using Deep Reinforcement Learning

This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot...

Filip Połatyński, Paweł Skruch · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.