Benign Nonconvex Landscape for Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action Spaces
We study the optimization landscape for infinite-horizon discounted Markov decision processes (MDPs) with general state and action spaces under structured stationary policy classes. A general weighted policy-iteration approach to establishing global convergence guarantees for policy gradient methods requires closure un...