Skip to content

BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services

Jul 2026 · arXiv.org · Vol abs/2607.04164 · 0 citations · 15 references
Computer Science

TL;DR

Inspired by the brownout paradigm in service computing, BrownoutMoE reorganizes experts into groups to improve utilization and system efficiency while maintaining service quality and introduces a grouping-consistent distillation process to produce deployable models that are compatible with standard inference pipelines.

Abstract

Mixture-of-Experts (MoE) large language models (LLMs) are increasingly deployed in Web-facing services, where inference must be both accurate and responsive under bursty demand. Although MoE models improve parameter efficiency through sparse expert activation, efficient MoE inference remains challenging in practice. A major reason is the highly imbalanced expert access pattern during inference: a few hot experts process most routed tokens, while many cold experts are rarely activated, leaving GPU parallelism underutilized. Existing systems mainly optimize runtime execution, such as scheduling, communication overlap, and kernel fusion, but usually preserve the original expert organization and therefore do not address the structural inefficiency caused by fragmented expert usage. In this paper, we present \textbf{BrownoutMoE}, a structure-aware optimization framework for efficient and accurate MoE inference services. Inspired by the brownout paradigm in service computing, BrownoutMoE reorganizes experts into groups to improve utilization and system efficiency while maintaining service quality. Specifically, we formulate layer-wise expert grouping as a learning problem and employ reinforcement learning to discover grouping strategies that minimize accuracy degradation. We further introduce a grouping-consistent distillation process to produce deployable models that are compatible with standard inference pipelines. Experimental results demonstrate that BrownoutMoE reduces accuracy degradation by up to 71.4% and improves throughput by up to 2.24x over baselines.

View source

Similar papers

Preprint Aug 2026

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...

Amjad Saab · 0 citations
Book Open access Sep 2026

MigMoE: Task-Aware Expert Migration for Faster and More Balanced Expert-Parallel MoE Inference

Sparse Mixture-of-Experts (MoE) architectures scale LLM capacity, but serving them with expert parallelism often suffers from the straggler effect caused by skewed token routing and uneven placement of hot experts across GPUs. Existing methods mitigate stragglers by adjusting token-to-expert distributions or by replica...

Xu Han, Zinuo Cai, Zhuo-Long Jiang et al. · 0 citations
Conference Open access Sep 2026

DoMoE: Domain-Aware Semantic Expert Prediction for Efficient MoE Inference Under Expert Offloading

DoMoE is proposed, a domain-aware MoE inference system that exploits domain locality in inference workloads and organizes routing information into domain-specific expert routing tables, restricts semantic matching to domain-relevant tokens and explicitly balances prediction accuracy against prediction overhead.

Yao Mu, Fa-Hao Chen, Wen-Bin Zhu et al. · 1 citation
#small language model Book Open access Aug 2026

Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism

EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.

Jia-Min Cao, Qingxu Li, Yaozhong Liu et al. · 2 citations · ⚡1
Preprint Aug 2026

SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

SAEM is proposed, a stage-aware MoE inference runtime that detects reasoning stage boundaries and exploits stage-level activation coherence to guide expert placement and achieves an average 1.33x throughput improvement over the strongest state-of-the-art caching and offloading baselines under constrained GPU memory.

Yujie Zhang, Bin Gao, Tulika Mitra · 0 citations
#large language models Book Open access Sep 2026

CARE-MoE: Correlation-Aware Expert Placement and Semantic Equivalence Routing for MoE LLM Inference on Edge Devices

CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.

Zhen-Yu Wang, Wei Li, Ao Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.