This work extends the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps under a response model in which each player conditions on the opponents'currently realized actions and evaluates continuation with that profile frozen.
Abstract
We study entropy-regularized exploratory control in finite $N$-player stochastic differential games under a response model in which each player conditions on the opponents'currently realized actions and evaluates continuation with that profile frozen. The resulting Gibbs best responses form a system of full conditional densities, which need not admit a common joint law. We characterize joint realizability by a cross-partial condition on the entropy-scaled Hamiltonian gradients and, on simply connected action domains, by an equivalent entropy-weighted potential structure. When compatibility fails, a coordinate-path construction yields a joint density whose full conditionals satisfy explicit quadratic Kullback--Leibler bounds. We extend the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps. A two-player linear-quadratic example illustrates the compatibility criterion and the associated learning procedure.
This work introduces an independently randomized formulation in which each stopping rule is represented by an adapted, nondecreasing cumulative stopping process, and identifies an exact-potential subclass with a closed-form threshold equilibrium.
We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space structure) and continuous/discrete action spaces. The setup covers many well-known applications in operations such as multi-product dynamic pric...
Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlated equilibrium, are computationally difficult, and thus, standard decentralized reinforcement-learning algorithms should not generally be expected to converge to them. In this paper, we study the equilibrium ge...
We study decentralized learning of Nash equilibria (NE) in infinite-horizon discounted Markov games under bandit feedback, focusing on Markov $\alpha$-potential games. We develop KL-projected natural policy gradient (NPG) algorithms in two settings: an episodic setting with frozen policies during sampling and a fully o...
A nonparametric distributional Bellman optimality operator for JMDPs is defined, and it is proved that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law.
Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al.· 0 citations
Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Finite-horizon zero-sum networked separable Markov games occupy an important middle ground: they retain global competition structure t...
Zai-Lin Ma· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.