#artificial intelligence
May 2026
Selective Off-Policy Reference Tuning with Plan Guidance
SORT adds a repair update for all-wrong prompts without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning.
A. Duc, Tien-Phat Nguyen, T. Nguyen et al.
· arXiv.org · 1 citation