Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute
This work trains from random initialization through searched self-play on a single eight-GPU node for 2.5 days, and investigates search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering.