On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion fo...
Yu-Hao Wang, Ruiyang Ren, Yi-Nan Zhang et al.· 0 citations
As Large Language Models (LLMs) evolve, parallel reasoning has emerged as a vital inference paradigm that enhances robustness by concurrently exploring multiple thought trajectories. Unlike fragile sequential methods, parallel reasoning expands inference breadth to significantly improve problem-solving performance. T...
Zi-Qi Wang, Bo-Ye Niu, Zi-Peng Gao et al.· National Science Review· 0 citations
FormalEvolve maintains a compilation-feasible archive for reuse and returns a deduplicated, semantically accepted repertoire for evaluation and downstream proving, and shows that archive-search gains persist with stronger seed and repair models.
Hai-Jian Lu, Wei Wang, Jing Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.