Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
This work extends the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps under a response model in which each player conditions on the opponents'currently realized actions and evaluates continuation with that profile frozen.