A novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head is introduced, enabling professional-level play on limited hardware where massive MCTS is infeasible.
Abstract
Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strategically fatal. This paper introduces a novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head. Unlike traditional value functions, the Belief head acts as an internal simulator and independent critic, modeling epistemic uncertainty and strategic stability. By integrating memory mechanisms (Transformer/GRU) to handle long-term dependencies and the Ko rule, and utilizing a gating mechanism to filter overconfident policy errors, our model shifts the burden of intelligence from runtime search to parametric"intuition."Experimental results demonstrate that this approach significantly improves search-free win rates and reduces hallucination, enabling professional-level play on limited hardware where massive MCTS is infeasible.
The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al.· 1 citation
Observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes, and these findings show that calibration analyses of repeated human–AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional...
Shreyan Biswas, Alexander Erlei, U. Gadiraju· Proceedings of the 2026 ACM...· 0 citations
An agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment is introduced, and it is shown that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying...
Jakub Rada, Viliam Lisý AI Center, Department of Computer Science et al.· 0 citations
Fast and accurate decisions are fundamental for adaptive behaviour. Theories of decision-making posit that evidence in favour of different choices is gradually accumulated until a critical value is reached. It remains unclear, however, which aspects of the neural code get updated during evidence accumulation. Here we i...
Dragan Rangelov, Sebastian Bitzer, Jason B. Mattingley· Journal of Neuroscience· 0 citations
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-...
Lucas Biechy, Cédric Eichler, Adrien Boiret et al.· 0 citations
A formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design of ERR+, a two-phase RLVR framework grounded in this observation.
Xinle Jiang, Min-Hao Wang, Wen Wu et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.