Continuous batching improves large language model (LLM) serving throughput, but long prompt prefills can delay decode iterations and violate inter-token latency objectives. Chunked prefill mitigates this interference, yet its chunk size is normally fixed: small chunks protect decode latency but repeatedly pay launch ov...
Si-Yu Song, Qi-Wei Bai, Jin-Bo Hao et al.· 1 citation
AnthroDial is presented, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment and shows that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
Wentao Liu, Si-Yu Song, Xi Chen et al.· arXiv.org· 0 citations
HeuristicEdu is presented, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO), and Scaffolding Effectiveness and Conversation Depth are introduced to evaluate outcomes beyond surface fluency.
Xiaokun Wang, Si-Yu Song, Wentao Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.