Skip to content
Preprint

Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching

Jul 2026 · 0 citations · 29 references
Mathematics

TL;DR

On-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data are developed, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data.

Abstract

This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.

View source

Similar papers

Preprint Aug 2026

Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control

Empirical results demonstrate that PI-VM matches SOTA precision with an order-of-magnitude efficiency gain in low-dimensional settings, while effectively mitigating mode collapse in high-dimensional scenarios, and offers a scalable solution for solving complex SOC problems.

Bangyan Liao, Cheng-Lei Yu, Yuchen Yang et al. · 0 citations
#reinforcement learning Open access Aug 2026

Decentralized strategies for finite population LQG social control: A reinforcement learning approach

This paper presents a novel model-free algorithm for the finite-population linear quadratic Gaussian (LQG) decentralized social control problem with multiplicative noise. The state and control weights in the cost functional are not limited to be positive semidefinite. For both finite-horizon and infinite-horizon cases,...

Liangyuan Guo, Bing-Chang Wang, Guangchen Wang · 0 citations
Preprint Aug 2026

Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints

A dual-ascent algorithm is proposed to solve the original problem as a constrained Markov decision process and it is proved that this formulation enjoys strong duality, thereby enabling it to reformulate the problem as an equivalent unconstrained one in the Lagrange dual framework.

Francesco Cordiano, Kang-Hui He, B. de Schutter · 0 citations
Preprint Aug 2026

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtai...

Xinyu Cao, Bing-Chang Wang, Ying Cao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.