Aug 2026· International Conference on Methods & Models in Automation & Robotics· pp. 303-308· 0 citations· 10 references
Abstract
This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot was developed using the Simscape Multibody toolbox, providing a physics-based simulation environment for training and evaluation. The proposed control approach employs a deep reinforcement learning agent trained using the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. The agent learns locomotion behaviors directly from interactions with the simulated environment, without relying on predefined gait trajectories or manually designed control laws. Through iterative training, the agent optimizes its policy to maximize a predefined reward function, enabling the robot to discover efficient and stable movement patterns. Simulation results demonstrate that the TD3-based approach is highly effective for continuous control tasks involving systems with complex nonlinear dynamics. The trained agent successfully learned locomotion strategies, including dynamic gaits with flight phases, highlighting the ability of reinforcement learning methods to handle naturally unstable behaviors that are difficult to design using classical control techniques.
Effective learning design guidelines for realizing arm-based locomotion on life-sized robotic hardware and expanding the traversable workspace of robots are provided.
Ayumu Iwata, Kento Kawaharazuka, Keita Yoneda et al.· 0 citations
. Quadrupedal robots exhibit strong mobility in complex environments where wheeled platforms often perform poorly, but their control remains difficult. In recent years, reinforcement learning (RL) has received growing attention in quadrupedal locomotion, as it supports direct policy optimization without relying entirel...
Chen Chang· Proceedings of the 3rd Inter...· 0 citations
This paper analyzes prevailing challenges and corresponding countermeasures regarding hardware deployment, sample efficiency, and model generalization capacity, and points out that further integration of multi-algorithms, optimization of sim-to-real transformation and overall strategy design will be the main trends in...
Context—Snake-like robots are biomimetic systems that can move effectively in narrow, complex, and restricted environments thanks to their modular and flexible body structures composed of numerous serially connected joints. These characteristics offer significant advantages, particularly in areas such as pipeline inspe...
Furkan Mezgil, M. Bingöl· Pamukkale Üniversitesi Mühen...· 0 citations
Reinforcement learning-based quadruped locomotion policies can exhibit command-tracking errors under terrain variations and unmodeled dynamics. This study proposes an online bounded Gaussian Process-enhanced model predictive control framework, termed Gaussian Process–Model Predictive Control–Reinforcement Learning(GP-M...