This work introduces an independently randomized formulation in which each stopping rule is represented by an adapted, nondecreasing cumulative stopping process, and identifies an exact-potential subclass with a closed-form threshold equilibrium.
Abstract
Finite-player nonzero-sum optimal stopping games typically lead to coupled equilibrium systems whose complexity grows rapidly with the number of players. We introduce an independently randomized formulation in which each stopping rule is represented by an adapted, nondecreasing cumulative stopping process. The canonical embedding preserves pure-profile payoffs, and a pure profile is a Nash equilibrium of the original game if and only if its embedding is a Nash equilibrium of the randomized game. We adopt the $\alpha$-potential approach to construct an $\alpha_N$-potential function, with the error $\alpha_N=O(N^{-1})$ under weak-interaction. We also identify an exact-potential subclass with a closed-form threshold equilibrium. For local stopped-status interactions, randomized payoffs admit a local stopped-mass representation, and potential maximization can be formulated as a multidimensional singular-control problem with local gradient constraints and a nonlocal condition for finite jumps. Under suitable regularity assumptions, we study the associated Hamilton-Jacobi-Bellman quasi-variational inequality and its regularity properties. For unknown model coefficients, we propose a bounded-intensity Potential-CT-DDPG learning algorithm. Numerical experiments closely match the analytical benchmark and yield estimated best-response improvements consistent with $N^{-1}$ scaling.
We study decentralized learning of Nash equilibria (NE) in infinite-horizon discounted Markov games under bandit feedback, focusing on Markov $\alpha$-potential games. We develop KL-projected natural policy gradient (NPG) algorithms in two settings: an episodic setting with frozen policies during sampling and a fully o...
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneously at every round $t...
We study feedback stabilization for the discrete-time system $x_{t+1}=f(x_t)+u_t+w_{t+1}$ in $\mathbb{R}^d$ with unknown $f$ and arbitrary bounded disturbances. For scalar plants, the sharp feedback capability threshold under generalized Lipschitz uncertainty is $3/2+\sqrt{2}$. We treat fully coupled vector-valued syst...
We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting...
Omar Abbadi, R. Laraki, P. Mertikopoulos· 4 citations· ⚡4
This paper studies $N$-player stochastic linear-quadratic (LQ) differential games from the perspective of $\alpha$-potential games. We first consider a closed-loop LQ game with multiplicative noise, where both the drift and the diffusion coefficients depend linearly on the state and the full control vector. For this mo...
Theoretical analysis and numerical experiments show the proposed methods substantially outperform the existing approaches to solve monotone linear-quadratic v-GNE problems.
Alberto Bemporad, T. Tatarenko· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.