In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify...
Salah Chikhi, Abdelhaq Chaoui, A. Ozdaglar et al.· 0 citations
This work introduces \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an EMA cascade for high-order optimism (ECHO), where EMA denotes exponential moving average, and leverages a new form of optimism inspired by modern filter design.
Mingyang Liu, Gabriele Farina, A. Ozdaglar· 5 citations· ⚡4
The proposed update incorporates the objective gradient inside the denoising step, yielding an inference-time method that uses only a pretrained denoiser and gradient evaluations and is analyzed as an inexact projected-gradient method for constrained optimization over learned feasible geometries.
R. Zhang, Jiawei Zhang, Gioele Zardini et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.