Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next mov...
Yutong Liang, Quanquan Peng, Matthew Kim et al.· 0 citations
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safet...
Bo-Yang Li, Matthew Kim, Sylvia L. Herbert· 0 citations
This work builds on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and proposes Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL.
Bo-Yan Li, Matthew Kim, Sylvia L. Herbert· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.