Sim-to-Real Reinforcement Learning for Ball-Balancing Locomotion on Quadruped Robots
Abstract
Non-prehensile manipulation of freely moving objects on a mobile base represents a significant challenge in underactuated robotics. This paper presents a sim-to-real reinforcement learning pipeline for a Unitree Go2 robot tasked with balancing a free-rolling ping-pong ball on a board mounted on its trunk while maintaining stable posture or tracking commanded velocities. Built on Legged Gym and the Genesis simulation, the proposed framework augments standard quadrupedal locomotion with ball-aware observations, task-specific reward design, and curriculum learning for progressively harder balancing and locomotion regimes. To improve transfer, the method incorporates domain randomization, camera-rate-compatible ball observations that mimic asynchronous visual feedback, and deployment-oriented safeguards such as smooth startup action blending. The learned policies are evaluated through a three-stage pipeline: large-scale training in Genesis, sim-to-sim validation in MuJoCo, and deployment on a physical Unitree Go2 using vision-estimated board-frame ball states. Experimental results show that the proposed framework can achieve both standing ball balance and ball-balancing locomotion on hardware, while additional comparisons between PPO and SAC highlight a trade-off between nominal task performance and disturbance robustness. These results suggest that reinforcement learning, when combined with transfer-aware observation design and deployment mechanisms, provides a practical approach for dynamic ball-balancing control on quadruped robots.