The Surprising Effectiveness of Approximate Value Iteration in Self-Play
The results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-lookahead greedy policies remain competitive with MCTS-based policies at substantially lower training and inference costs.
Raphaël Boige, Amine M. Boumaza, Bruno Scherrer
· 0 citations