This project implements AC4MPC, in which a reinforcement-learning agent is trained on the exact cost a model predictive controller optimizes, and its learned value function (critic) is embedded in the MPC as a terminal cost.
Model predictive control (MPC) provides interpretable, tunable locomotion controllers grounded in physical models, but its robustness depends on frequent replanning and is limited by model mismatch and real-time computational constraints. Reinforcement learning (RL), by contrast, can produce highly robust behaviors thr...
Seungmin Jeon, Ho Jae Lee, Seung-Woo Hong et al.· IEEE Transactions on robotic...· 8 citations
Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy opti...
Pengqin Wang, Qi-Ming Zhang, Shao-Jie Shen et al.· 0 citations
A training method for HIL online reinforcement learning for real robots that automatically switches between learning from interventions and on-policy self-improvement, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift.
This work formalizes the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and shows that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate.
Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba· 0 citations
This work proposes Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone, and introduces a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalizatio...
Integrated deep reinforcement learning (DRL) and model predictive control (MPC) methods are increasingly used to control autonomous systems by combining their complementary capabilities. DRL learns control policies through interaction with the environment. MPC uses a system model to optimize control inputs while accoun...
Giray Önür, A. Dabiri, B. de Schutter· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.