Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games
Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Finite-horizon zero-sum networked separable Markov games occupy an important middle ground: they retain global competition structure t...