On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...
Bing Shao, Jia-Zheng Zhang, Long Ma et al.· 5 citations
Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...
Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al.· 0 citations
A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.
Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al.· arXiv.org· 0 citations
CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Bo-Yang Liu, Sen-Jie Jin, Pei-Xin Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.