Jul 2026· Proceedings of the 3rd Foundations of Process/Product Analytics and Machine Learning (FOPAM 2026)· pp. 12-12· 0 citations· 1 references
TL;DR
This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality.
Abstract
Optimal decision-making under uncertainty is a shared challenge across modern chemical, manufacturing, and energy systems that increasingly demand safe, data-driven autonomy. This talk revisits optimal control through the lens of the Bellman equation, emphasizing how optimal control theory and reinforcement learning have developed complementary, yet largely disconnected, perspectives on global optimality. In one view, central to reinforcement learning, the Bellman equation defines a global optimality condition that guides iterative policy learning from interacting with the system, but typically yields opaque control laws that are difficult to interpret, and deploy in safety-critical settings. In another view, widely adopted in model predictive control (MPC), the Bellman equation underpins tractable finite-horizon optimizations that deliver interpretable, constraint-aware, and modular local controllers, yet without explicit guarantees on alignment with global optimality. Building on these ideas, we introduce a local–global paradigm that treats MPC and related optimization-based controllers as structured function approximators designed to approximately satisfy the global Bellman optimality condition. We discuss algorithmic strategies for learning interpretable local decision makers whose adaptation is guided by Bellman residuals, along with the benefits and practical challenges that arise in terms of stability, constraint satisfaction, and sample efficiency. These concepts are illustrated through case studies that unify reinforcement learning and MPC for safe, high-performance control in complex, uncertain dynamical systems. The talk concludes by outlining open problems and research opportunities in learning interpretable control policies that achieve globally optimal performance while retaining the transparency and reliability required for real-world process control and optimization applications.
This work separates local mixed stability, an occupancy-weighted first-order relaxation that is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies, and extends both notions to n-player performative Markov games, obtaining local stability with no assumption on t...
A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.
This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level, achieving higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen...
Seungmin Jeon, Ho Jae Lee, Seung-Woo Hong et al.· IEEE Transactions on robotic...· 10 citations
Geometric Distributional Control is validated on structured multilevel optimization and SUMO route-progress driving, where it improves over known-only solvers and learning baselines while preserving scaffold-enforced feasibility.
Comparative control theory analysis reveals a fundamental trade-off: transitioning from continuous to discrete rule-based control induces high-frequency Bang-Bang actuation and a stable bimodal limit cycle.
This work proposes Transition Occupancy Matching as a unifying principle to resolve policy and dynamics shifts within a single mathematical framework and introduces Occupancy-Matching Policy Optimization (OMPO), a novel algorithm that optimizes a surrogate objective explicitly correcting for transition discrepancies.
Yu Luo, Lei Lv, Fu-Chun Sun et al.· National Science Review· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.