Hyperparameter Scaling Laws Across MoE Sparsity
This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.
Chang-Xin Tian, Kun-Long Chen, Jia Liu et al.
· 0 citations