Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...
Hao-Ran Zhao, Wei Du, Dingwen Yang et al.· 0 citations
On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...
Bing Shao, Jia-Zheng Zhang, Long Ma et al.· 5 citations
Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
EntangleCodec is proposed, a unified discrete audio tokenizer that learns caption-aligned semantic-acoustic representations before quantization that achieves reconstruction quality competitive with specialized codecs, outperforms all codec-based baselines on audio understanding, and supports both TTS and TTA generation...