Skip to content
Open access

Online GP-MPC Command Supervision for Robust Reinforcement Learning-Based Quadruped Locomotion

Sep 2026 · Biomimetics · Vol 11 · 0 citations · 25 references
Medicine

Abstract

Reinforcement learning-based quadruped locomotion policies can exhibit command-tracking errors under terrain variations and unmodeled dynamics. This study proposes an online bounded Gaussian Process-enhanced model predictive control framework, termed Gaussian Process–Model Predictive Control–Reinforcement Learning(GP-MPC-RL), for command-level supervision of a pretrained locomotion policy. A frozen PPO policy generates the low-level locomotion behavior, while an acados-based MPC supervisor adjusts the velocity command using a nominal command-response model. An online Gaussian Process learns the one-step residual between the nominal prediction and measured robot response, and its uncertainty-weighted forward-velocity correction is incorporated into the MPC prediction. The framework was evaluated in Isaac Lab using a Unitree Go2 quadruped robot model over 20 paired rough-terrain trials at target velocities of 0.3, 0.5, and 0.7 m/s; GP-MPC-RL reduced the mean forward-velocity root mean square error (RMSE) relative to PPO by 38.3%, 21.3%, and 10.1%, respectively. Under a 5 kg payload, GP-MPC-RL reduced velocity RMSE by 29.0% relative to MPC-RL and reduced the 0–5 kg payload-induced degradation by 53.0% (p = 0.019). The average supervisor computation time was 0.112 ms. These results indicate that GP residual adaptation is particularly effective when the nominal command-response model becomes inaccurate, improving robustness without retraining the underlying locomotion policy.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.