Skip to content
Preprint

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

Aug 2026 · 0 citations · 26 references
Mathematics

TL;DR

A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain.

Abstract

Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.

View source

Similar papers

Sep 2026

Carleman approximation based adaptive optimal control design of nonlinear systems: A three-phase policy iteration approach.

As one efficient algorithm in adaptive dynamic programming to address adaptive optimal control problems, policy iteration always involves an initial admissible control guess during iteration. Such an initialization process is nontrivial, especially for unstable systems or when system dynamics are totally unknown. To ci...

Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al. · 0 citations
Open access Sep 2026

Reinforcement-Learning Robust Tracking Control of Discrete-Time Systems with Diagonal Scaling

This paper addresses a robust tracking problem for linear discrete-time systems by proposing a reinforcement learning (RL) control method based on a diagonal-scaling strategy, offering a solution tailored to the demands of enhanced reliability. To overcome a common limitation in policy-iteration-based RL design, namely...

Kan-Yang Jiang, Zheng Gao, Ye Zeng et al. · 0 citations
Preprint Aug 2026

Policy Iteration for Linear-Quadratic Stochastic Differential Games with State- and Control-Dependent Noise

This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fr\'echet derivative of the sequential PI map...

Karl Handwerker, Felix Thömmes, Lucas Günther et al. · 1 citation
Preprint Sep 2026

Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers

Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and...

A. Nguyen, Leilei Cui · 0 citations
Sep 2026

Event-triggered reinforcement learning-based safe control for stochastic systems subject to asymmetric input constraints and unknown dynamics.

This paper investigates the safe optimal control (SOC) for input-constrained unknown stochastic systems via adaptive dynamic programming (ADP) and generalized fuzzy hyperbolic model (GFHM). Firstly, a GFHM is employed to approximate the unknown nonlinear terms of the stochastic system, thereby eliminating the need for...

Yu-Ling Liang, Feng-Lin Qin, Lei Liu et al. · 0 citations
Preprint Aug 2026

Iterative State- and Control-Dependent Model Predictive Control: A Jacobian-Free Formulation for Constrained Nonlinear Systems

This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex o...

Mohammadreza Kamaldar · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.