Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-s...
Kai-Rong Luo, Jia-Rui Cui, Yao-Rui Yin et al.· 0 citations
Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.
Shaowen Wang, Ge Zhang, Kai-Rong Luo et al.· 4 citations
This report presents an open pretraining recipe that trains a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs, and derives a Puro Cost Scaling Law that relates training cost to average model performance.
Kairong Luo, Jia-Rui Cui, Yao-Rui Yin et al.· 0 citations
This paper introduces RoPE-based Block-wise Sparse Attention (RoBSA), a method designed specifically for MLA during the decoding stage of model inference that significantly reduces end-to-end inference latency in the decoding stage by up to 2 .
Xinyu Shi, Kairong Luo, Zhen Zheng et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.