BOW: Training Language Models to Reason Over Plausible Next Words
BOW is introduced, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space, and human evaluation shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these trajectories remain predictive.