Reinforcement Learning for Bipedal Locomotion Using Minimal Instrumentation with a Single Inertial Measurement Unit
Abstract
This work presents a simulation-based validation framework for locomotion control on a custom-built 13 DoF bipedal robot using only signals derivable from a 6-axis inertial measurement unit (3D angular velocity and 3D gravity vector projection) as actor observations. The system employs the Genesis World simulator and the rsl-rl-lib library with PPO and a privileged critic architecture, where the actor only accesses IMU signals, reference speed commands, and past actions, without joint encoders or additional exteroceptive sensors. Training converges in 4,000 iterations (393M environment steps with 4,096 parallel environments): mean reward scales from 1.82 to 118.19, mean episode length goes from 22.8 to 1,007.6 steps, and success rate reaches 66.3%, with 99.8% vertical stability. The automatic curriculum unlocks running gait at iteration 126. This minimal instrumentation approach reduces hardware costs and facilitates replication in engineering laboratories with limited budgets.