Preprint
Jul 2026
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
BPO is instantiate as Branching Policy Optimization (BPO), a sandbox-native RL algorithm that adaptively snapshots the sandbox at high-entropy decision points along a backbone trajectory, and proves this estimator is unbiased and has strictly lower variance than the trajectory-level baseline, with the reduction equal to the prefix-explained portion of return variance.
Bowei He, Yankai Chen, Xiaokun Zhang et al.
· 1 citation