#natural language process...
Jun 2026
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development, is introduced, which outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.
Byung-Kwan Lee, Ximing Lu, Shi-Zhe Diao et al.
· arXiv.org · 2 citations