Jul 2026
Provably Optimal Learning Algorithms for Assistance Games
The notion of assistance regret is introduced: the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs, is introduced.
Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan et al.
· arXiv.org · 0 citations