Jul 2026
Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes
This paper provides the first finite-time convergence guarantees for this algorithm in this setting, for which it is proved that NPG converges sublinearly with a rate of $\mathcal{O}(H^{2}/t)$ after $t$ iterations, where $H$ is the horizon length.
Asha Barua, S. Khodadadian
· arXiv.org · 0 citations