Uncertainty-Budgeted Rollout Allocation for Self-Improving Reasoning Models
Abstract
The pursuit of artificial general intelligence has increasingly focused on enhancing the reasoning capabilities of large language models through self-improvement mechanisms. Central to these mechanisms is the generation of rollouts, which are simulated reasoning paths that allow models to explore multiple solution trajectories before committing to an final answer. However, generating an exhaustive set of rollouts for every query is computationally prohibitive and often inefficient, as models frequently allocate equivalent computational resources to trivial problems and highly complex tasks. This paper proposes a novel framework termed Uncertainty-Budgeted Rollout Allocation, which dynamically distributes computational resources based on the inherent epistemic and aleatoric uncertainty detected during the reasoning process. By mathematically formulating the rollout allocation problem as an uncertainty-constrained optimization task, the proposed methodology directs generation budgets toward reasoning nodes that exhibit high variance in predictive confidence, thereby maximizing the informative value of each simulated trajectory. Extensive empirical evaluations across diverse logical and mathematical reasoning benchmarks demonstrate that the proposed allocation strategy significantly improves reasoning accuracy while reducing overall computational overhead. The results reveal that dynamically budgeting rollouts based on real-time uncertainty metrics not only accelerates the self-improvement cycle but also mitigates the risk of reward hacking and premature convergence in reasoning models.