TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.