Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) a...
Wei-Li Xu, Ji-Sen Li, Yu-Qing Jian et al.· 0 citations
Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.
Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar et al.· 1 citation
Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models...
Fengxiang Bie, Yu-Qing Jian, Yi-Fan Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.