A three-player game is constructed with a preferentially stable set whose span is dynamically unstable, showing that preferences do not suffice as a criterion of dynamic stability and bridges the gap via the notion of resilience under aggregate deviations.
Abstract
We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players'learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames -- i.e. subsets of pure profiles obtained by restricting players'action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.
Rank-one games already separate efficient equilibrium computation from strategic planning against a learning opponent, and this result shows that this equilibrium tractability does not extend to planning against learning dynamics.
This work proposes a lightweight two-feature embedding that captures fundamental behavioural demands: the entropy of the Nash equilibrium and the sensitivity of optimal responses to an opponent's action, and shows that this embedding reliably predicts performance changes on held-out games.
Joshua Caiata, Sreepriya Pulyassary, Xiang Li et al.· arXiv.org· 0 citations
This paper analyzes the play that is generated when each agent follows a strategy that optimizes against an adversary, and considers the two known explicit constructions of optimal strategies.
Shaull Almagor, Guy Avni, Julian Ewaied· 0 citations
We study a decision-maker who explores --- dynamically choosing what to learn --- before stopping to act. We first reduce this dynamic control problem to a static one: any exploration-and-stopping strategy is equivalent to a choice of the joint distribution of the stopped state and the stopping time, subject to one inf...
This work extends the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps under a response model in which each player conditions on the opponents'currently realized actions and evaluates continuation with that profile frozen.
This work comprehensively investigates the concept of constant-memory strategies in stochastic games, and uncovers the connection between decision models in single-agent and multi-agent contexts.
Feng-Ming Zhu, Fang-Zhen Lin· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.