Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

BPO is instantiate as Branching Policy Optimization (BPO), a sandbox-native RL algorithm that adaptively snapshots the sandbox at high-entropy decision points along a backbone trajectory, and proves this estimator is unbiased and has strictly lower variance than the trajectory-level baseline, with the reduction equal to the prefix-explained portion of return variance.

Bowei He, Yankai Chen, Xiaokun Zhang et al. · 1 citation