Skip to content

Author

Kaifeng Lyu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Aug 2026

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

This report presents an open pretraining recipe that trains a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs, and derives a Puro Cost Scaling Law that relates training cost to average model performance.

Kairong Luo, Jia-Rui Cui, Yao-Rui Yin et al. · 0 citations
Preprint Jul 2026

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D

2D-RoPE is introduced, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID, and suggests that viewing text in 2D can benefit language modeling.

Haodong Wen, Yiran Zhang, Yingfa Chen et al. · 0 citations