Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
This work formalizes fork placement as locating the pivots of the chain's value curve, where the expected outcome turns, and proposes belief-shift branching, which read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most.
Bin Lei, Yu Li, Prafulla Kumar Choubey et al.
· 0 citations