Skip to content

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

Sep 2026 · 0 citations · 54 references
Computer Science Engineering

TL;DR

Solver-Gradient Guided Reinforcement Learning is proposed, a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation that reaches PPO's best closed-loop return with up to 70.6% fewer samples, and outperforms GB-PL baselines by at least 54% in closed-loop return.

Abstract

In Model Predictive Control (MPC), cost-function weights shape closed-loop behavior, yet changing conditions often make fixed parametrizations suboptimal and motivate context-dependent online adaptation. Learning such policies is difficult because behavior depends implicitly on numerical MPC solutions, producing nonlinear, potentially nonsmooth, long-horizon dependencies on policy parameters. This creates a bias-variance tradeoff: Reinforcement Learning (RL) optimizes realized closed-loop return from environment samples but is sample-inefficient, whereas Gradient-Based Policy Learning (GB-PL) uses low-variance solver gradients from differentiable MPC to optimize surrogate losses on predicted trajectories but can be biased under model mismatch. We propose Solver-Gradient Guided Reinforcement Learning (SG-RL), a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation. SG-RL keeps sampled closed-loop return as the objective and uses bounded solver-derived gradients as auxiliary guidance to improve stability and sample efficiency. We instantiate SG-RL in Proximal Policy Optimization (PPO) with four modular algorithms that inject solver-gradient guidance into actor-update scaling, policy loss, advantage estimation, and value-function learning. On two full-scale autonomous racing platforms with intentional model mismatch, SG-RL reaches PPO's best closed-loop return with up to 70.6% fewer samples, outperforms GB-PL baselines by at least 54% in closed-loop return, and generalizes zero-shot to unseen environments.

View source

Similar papers

#machine learning Preprint Sep 2026

ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards

Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias throu...

Hsiao-Ru Pan, Florent Draye, Bernhard Scholkopf · 0 citations
#reinforcement learning Preprint Aug 2026

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.

Hossein Abdi, S. Dash, Ming-Fei Sun · 0 citations
Open access Oct 2025

Residual MPC: Blending Reinforcement Learning With GPU-Parallelized Model Predictive Control

Model predictive control (MPC) provides interpretable, tunable locomotion controllers grounded in physical models, but its robustness depends on frequent replanning and is limited by model mismatch and real-time computational constraints. Reinforcement learning (RL), by contrast, can produce highly robust behaviors thr...

Seungmin Jeon, Ho Jae Lee, Seung-Woo Hong et al. · 8 citations
Preprint Aug 2026

Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics

A model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law is proposed, ensuring that each model operates within its reliability regime.

A. Maalberg, A. Neumann, J. Knobloch · 0 citations
#artificial intelligence Preprint Sep 2026

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead imp...

Hyukjun Yang, Jongchan Park, Narim Jeong et al. · 0 citations
Open access 2026

GCR-RL: Gradient Control Reward Shaping for Reinforcement Learning

This work introduces Gradient Control Rewards (GCR), an interpretable, control-inspired reward-design methodology for accelerating agent training by modulating the reward signal based on the temporal dynamics of system error, Inspired by classical control theory.

Anas Aburaya, H. Selamat, M. Muslim et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.