Skip to content
Open access

Local-Global Learning of Interpretable Control Polices: The Interface between MPC and Reinforcement Learning

Jul 2026 · Proceedings of the 3rd Foundations of Process/Product Analytics and Machine Learning (FOPAM 2026) · pp. 12-12 · 0 citations · 1 references

TL;DR

This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality.

Abstract

Optimal decision-making under uncertainty is a shared challenge across modern chemical, manufacturing, and energy systems that increasingly demand safe, data-driven autonomy. This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality. In one view, central to reinforcement learning, the Bellman equation defines a global optimality condition that guides iterative policy learning from interacting with the system, but typically yields opaque control laws that are difficult to interpret, and deploy in safety-critical settings. In another view, widely adopted in model predictive control (MPC), the Bellman equation underpins tractable finite-horizon optimizations that deliver interpretable, constraint-aware, and modular local controllers, yet without explicit guarantees on alignment with global optimality. Building on these ideas, we introduce a local–global paradigm that treats MPC and related optimization-based controllers as structured function approximators designed to approximately satisfy the global Bellman optimality condition. We discuss algorithmic strategies for learning interpretable local decision makers whose adaptation is guided by Bellman residuals, along with the benefits and practical challenges that arise in terms of stability, constraint satisfaction, and sample efficiency. These concepts are illustrated through case studies that unify reinforcement learning and MPC for safe, high-performance control in complex, uncertain dynamical systems. The talk concludes by outlining open problems and research opportunities in learning interpretable control policies that achieve globally optimal performance while retaining the transparency and reliability required for real-world process control and optimization applications.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Local and Global Stability in Performative Reinforcement Learning

This work separates local mixed stability, an occupancy-weighted first-order relaxation that is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies, and extends both notions to n-player performative Markov games, obtaining local stability with no assumption on t...

Debmalya Mandal · 0 citations
#reinforcement learning Preprint Aug 2026

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.

Hossein Abdi, S. Dash, Ming-Fei Sun · 0 citations
Open access Oct 2025

Residual MPC: Blending Reinforcement Learning With GPU-Parallelized Model Predictive Control

This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level, achieving higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen...

Seungmin Jeon, Ho Jae Lee, Seung-Woo Hong et al. · 10 citations
Jul 2026

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

Comparative control theory analysis reveals a fundamental trade-off: transitioning from continuous to discrete rule-based control induces high-frequency Bang-Bang actuation and a stable bimodal limit cycle.

Shaker Al-Tamari, Waled Kadour · 0 citations
Open access Aug 2026

Embodied Learning under Policy and Dynamics Shifts

This work proposes Transition Occupancy Matching as a unifying principle to resolve policy and dynamics shifts within a single mathematical framework and introduces Occupancy-Matching Policy Optimization (OMPO), a novel algorithm that optimizes a surrogate objective explicitly correcting for transition discrepancies.

Yu Luo, Lei Lv, Fu-Chun Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.