Preprint
Aug 2026
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
A compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons is proposed.
Nayeon Kim, Hojin Lee, Yunju Bak et al.
· 0 citations