Skip to content
Preprint

(Cheap) Stochastic Policy Gradient Converges with High Probability for Linear Quadratic Regulator

Sep 2026 · 0 citations · 42 references
Mathematics

Abstract

We study the convergence of the vanilla stochastic policy gradient method applied to the linear quadratic regulator (LQR) problem. The method is cheap in the following sense: (1) at each iteration only $\tilde{O}(1)$ interactions with the environment are needed, therefore allowing frequent policy improvement steps, and (2) to ensure stability throughout and convergence to an $\epsilon$-optimal policy with probability $1-\delta$, only $O(\mathtt{Polylog}(1/\delta)/\epsilon)$ interactions are needed. To the best of our knowledge, this appears to be the first time that a stochastic model-free policy optimization method for LQR converges with high probability with $\tilde{O}(1)$ per-iteration computation and polylogarithmic dependence on the confidence level. The convergence analysis presented here is agnostic to LQR specifics and hence could be potentially generalized to a broader class of problems.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.