Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling...
Zhi-Heng Hu, Yi-Xun Wei, Jian Zhou et al.· 0 citations
This work introduces Farseer, a novel and refined scaling law offering enhanced predictive accuracy across scales, and provides new insights into optimal compute allocation, better reflecting the nuanced demands of modern LLM training.
Houyi Li, Wen-Zheng Zheng, Qiufeng Wang et al.· Neural Information Processin...· 4 citations· ⚡1
MISA-T, a routing-layer admission policy for mixed rollout serving that combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting, is presented.
Zetao Hong, Song Yuan, Yuanhao Ding et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.