Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) a...
Wei-Li Xu, Ji-Sen Li, Yu-Qing Jian et al.· 0 citations
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable latent linear RNN g...
Zi-Yan Chen, Zhong-Zhu Zhou, Pei-Lin Liu et al.· 0 citations
Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.
Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar et al.· 1 citation
Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models...
Fengxiang Bie, Yu-Qing Jian, Yi-Fan Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.