Skip to content

The Surprising Effectiveness of Approximate Value Iteration in Self-Play

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

The results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-lookahead greedy policies remain competitive with MCTS-based policies at substantially lower training and inference costs.

Abstract

Combining search with function approximation has driven major advances in game-playing programs, making self-play algorithms more competitive than ever. Still, the computational overhead of the most popular methods, based on Monte Carlo Tree Search (MCTS), can be substantial. In this work, we investigate whether simpler methods remain competitive in non-trivial, moderately sized games such as Connect Four, Hex(7x7) and synthetic games. We train a minimal self-play implementation of Approximate Value Iteration (AVI) and use ground-truth oracles for exact evaluation. Contrary to expectations, our results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-lookahead greedy policies remain competitive with MCTS-based policies at substantially lower training and inference costs. Preliminary experiments on Othello and Go(9x9) show that AVI trains stably on larger games and learns effective value functions. These findings suggest that the success of MCTS-based methods may have eclipsed simpler approaches that have become increasingly practical with modern deep-learning tools.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute

This work trains from random initialization through searched self-play on a single eight-GPU node for 2.5 days, and investigates search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering.

Bertil Braun · 0 citations
#artificial intelligence Conference Sep 2026

Online Robust Reinforcement Learning Through Monte-Carlo Planning

A new robust variant of MCTS that mitigates dynamical model ambiguities to bridge the gap between simulation-based planning and real-world deployment and empirical evidence is provided that this method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distributio...

T. Dam, Kishan Panaganti, Brahim Driss et al. · 4 citations
#machine learning Preprint Sep 2026

Constant regret in general games via higher-order optimism

We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting...

Omar Abbadi, R. Laraki, P. Mertikopoulos · 4 citations · ⚡4
Preprint Sep 2026

Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequ...

Ondrej Kubícek, Viliam Lisý, Tuomas Sandholm · 0 citations
#machine learning Preprint Sep 2026

CompassPlay: Rewarding the Proposer for Where It Moves the Solver

In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks wh...

Sophia Xiao Pu, Xi-Meng Sun, Jiang Liu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.