Skip to content

Hyperparameter Scaling Laws Across MoE Sparsity

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.

Abstract

Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.

View source

Similar papers

#machine learning Preprint Sep 2026

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of...

Atindra Jha, Margaret Li, J. Leskovec et al. · 1 citation
#machine learning Preprint Sep 2026

Scaling Zero-Order Pretraining through Model Sharding

Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient variance grows with perturbed dimension, inhibiting large-model training. Sharded Optimization Mixture of Assemblies (SOMA) trains LSTM experts independently on $N$ data...

Francois Chaubard, M. Kochenderfer, Chris Ré · 0 citations
Preprint Aug 2026

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

MOSAIC is developed, which formulates model architecture and systems co-design as an optimization problem that couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout.

Soumajyoti Sarkar, Yu-Xin Tang, Sheng Zha · 0 citations
#artificial intelligence Preprint Aug 2026

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

This study establishes a first baseline and scaling procedure for the development of future OpenEuroLLM models, and characterize the dependence of loss on model capacity and dataset size, evaluating recently proposed scaling forms that explicitly model their interaction.

Niccolò Ajroldi, Diana-Alexandra Onutu, Haider Al-Tahan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

This work introduces Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates optimal hyperparameter scaling laws through adaptive experimentation using less than one-tenth of the computational budget required by convention...

Zhiliang Chen, S. Ament, David Eriksson et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.