GAD-RL is introduced, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions, and weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that...
Bao-De Wang, Zu-Ming Huang, Ke Ren et al.· 0 citations
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the...
Wei Chu, Yuan-Zhe Dong, Ke Tan et al.· 0 citations
This work presents Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora.
Zuming Huang, Jun Huang, Kexuan Ren et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.