From Base Rollouts to RL Reasoning: A Budgeted Search Perspective
Results support a qualified internalized-search reading: under the recipe the authors test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search.
Wen-He Sun, Cunxiang Wang, Zijun Yao et al.
· 0 citations