Skip to content
Preprint

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

Aug 2026 · 1 citation · 39 references
Computer Science

TL;DR

A deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss that employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations.

Abstract

Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.

View source

Similar papers

Open access Sep 2026

Extending the Speed Limit of Quadrupedal Locomotion via Refined Actuator Modeling and Adaptive Command Scheduling

Achieving high-speed locomotion in quadrupedal robots remains highly challenging, as actuators operate near their physical limits and exhibit pronounced nonlinearities. However, many existing methods neglect actuator nonlinearities and physical constraints during training, leading to a significant sim-to-real gap under highly dynamic motions and limiting achievable performance. To address this issue, we propose a high-speed locomotion framework that reduces sim-to-real discrepancies and stabilizes learning over a wide command distribution. A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque-speed envelope. In addition, a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training. Experiments on the 36.5 kg quadruped BlackPanther2 (BP2) demonstrate speeds of up to 13.2 m/s on a treadmill and 11.65 m/s outdoors, establishing a new state-of-the-art and, to the best of our knowledge, a world record for quadrupedal robot locomotion. The results further highlight the importance of accurate actuator modeling in preventing non-physical policy exploitation, and show that ACS improves robustness without sacrificing performance.

Yu-Cheng Tao, Shao-wen Cheng, Guo-Rong Lan et al. · 0 citations
Preprint Aug 2026

Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space

This work proposes a hierarchical reinforcement learning pipeline that empowers the robots to perform aggressive locomotion through constrained obstacles--a narrow gate, extending the lifelike agility of legged robots to match that of their biological counterparts.

Zeren Luo, Jiahui Zhang, Yimin Han et al. · 1 citation
Jul 2026

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured>50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.

Javier C. Weddington, Bence P. Ölveczky, S. Baccus · 0 citations
Conference Aug 2026

Reinforcement Learning for Bipedal Locomotion Using Minimal Instrumentation with a Single Inertial Measurement Unit

This work presents a simulation-based validation framework for locomotion control on a custom-built 13 DoF bipedal robot using only signals derivable from a 6-axis inertial measurement unit (3D angular velocity and 3D gravity vector projection) as actor observations. The system employs the Genesis World simulator and the rsl-rl-lib library with PPO and a privileged critic architecture, where the actor only accesses IMU signals, reference speed commands, and past actions, without joint encoders or additional exteroceptive sensors. Training converges in 4,000 iterations (393M environment steps with 4,096 parallel environments): mean reward scales from 1.82 to 118.19, mean episode length goes from 22.8 to 1,007.6 steps, and success rate reaches 66.3%, with 99.8% vertical stability. The automatic curriculum unlocks running gait at iteration 126. This minimal instrumentation approach reduces hardware costs and facilitates replication in engineering laboratories with limited budgets.

Juan Esteban Gomez Lopez, Yesid Eugenio Santafe Ramon · 0 citations
Preprint Aug 2026

Rapid Embodiment Adaptation for Quadrupedal Locomotion

Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies often break when hardware properties shift. We introduce an online embodiment adaptation framework for quadrupedal locomotion that infers embodiment parameters from short interaction histories and conditions control on the inferred hardware state. Our method pairs a generalist policy trained under embodiment randomization with a lightweight adaptation module that identifies physical changes within half a second. We evaluate two representative forms of embodiment variation: joint-range constraints and trunk-mass changes, corresponding to joint-level kinematic degradation and body-level dynamic variation. In simulation, the module accurately estimates these changes and enables closed-loop control that substantially outperforms policies conditioned directly on interaction history. On a real Unitree Go2 robot, our system maintains stable locomotion under severe instances of the evaluated changes, including a fully locked leg and a 5 kg payload, where non-adaptive methods fail. These results demonstrate the practicality of explicit online embodiment identification for rapid adaptation to joint-limit and payload-mass changes, and provide a step toward handling broader forms of uncertain, degraded, or changing robot hardware.

Di-Chen Li, Bo Ai, Nico Bohlinger et al. · 1 citation
Open access Aug 2026

A unified CPG-based and multi-agent control framework for low-cost quadruped robots

Experimental results demonstrate that the integrated system improves locomotion stability, energy efficiency, and terrain adaptability compared with baseline controllers, highlighting the effectiveness of combining a structured gait prior, lightweight residual coordination, and hardware-aware deployment for practical quadruped locomotion.

Li-Kai Wu, Mei-Na Zhang, Wen-Zhe Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.