Skip to content
Preprint

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

Aug 2026 · 0 citations
Computer Science

TL;DR

MOSAIC is developed, which formulates model architecture and systems co-design as an optimization problem that couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout.

Abstract

In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate systems stage then optimizes the implementation for hardware efficiency. In this work, we develop MOSAIC, which formulates model architecture and systems co-design as an optimization problem. MOSAIC couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout. We instantiate the framework for sparse Mixture-of-Experts (MoE) language models, where expert count, routing sparsity, and other MoE layer dimensions affect both the loss and systems efficiency. We fit a scaling law on sparse MoE models trained on text data, whose scaling dimensions include the sparsity factor, which is the fraction of model parameters inactive per token in a forward pass. The scaling law sweeps in our work span active parameters from $104$ million to $2.7$ billion and total model sizes reaching $79$ billion parameters. We show that, within the calibrated sparsity range, an efficiency-agnostic model-FLOPs budget admits no interior optimal sparsity. The fitted loss decreases monotonically with sparser models and the compute optimum lies at the upper boundary of the data support. An optimal sparsity in MoE models instead emerges under the cluster's systems constraints, as captured by MOSAIC. Our results argue for a shift towards unified architecture and systems co-design for frontier language model training.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Hyperparameter Scaling Laws Across MoE Sparsity

This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.

Chang-Xin Tian, Kun-Long Chen, Jia Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Towards a Statistical Understanding of Mixture-of-Experts

This paper derives oracle risk bounds for learning dense and sparse routing with evolving experts, separating approximation, expert-learning, and router-estimation errors, and characterize how sparse Top-K routing can retain the benefits of localized aggregation while controlling per-input computation.

Si-Yuan He, Bo-Kai Yang, Jie Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be suff...

Dohyeon Kim, Bedionita Soro, Sung Ju Hwang · 0 citations
Preprint Aug 2026

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is proposed, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias and provides a tunable accuracy-expert-activation trade-off on DeepSeek-MoE-16B-Chat.

Rong-Feng Wang, Shitao Weng, Zhiquan Wang et al. · 0 citations
Preprint Aug 2026

Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANs

A multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively, significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in...

Xiaowen Cao, Zhonghao Lyu, Shicheng Chu et al. · 0 citations
#small language model Preprint Aug 2026

Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections

This work introduces Mixture of Channel Experts (MoCE), a structured sparse channel-mixing layer, inspired by MoE, that replaces pointwise (1x1) channel-reduction projections and matches or exceeds dense baselines and prior channel-selection methods while reducing MACs by 16.7% and end-to-end latency.

Elian Iluk, Gil Ben-Artzi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.