Skip to content

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

Sep 2026 · 0 citations · 56 references
Computer Science

TL;DR

OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs by reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit.

Abstract

Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE{\dag}, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.

View source

Similar papers

Preprint Oct 2026

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization...

Damiano Marsili, Raphi Kang, Aditya Mehta et al. · 0 citations
#machine learning Preprint Sep 2026

Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ran...

Zheng Lin, Shao-Ke Fang, Yu-Xin Zhang et al. · 0 citations
Preprint Aug 2026

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

MetaNet is proposed, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias and provides a tunable accuracy-expert-activation trade-off on DeepSeek-MoE-16B-Chat.

Rong-Feng Wang, Shitao Weng, Zhiquan Wang et al. · 0 citations
#machine learning Preprint Aug 2026

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This...

Si-Cheng Xu, Hao Shi, Wei Zhang et al. · 0 citations
Preprint Aug 2026

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

ExFold is proposed, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode that calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts.

Jun-Tong Wu, Yi-Fei Liu, Jun-Yi Chen et al. · 2 citations · ⚡2
Preprint Aug 2026

Shape Mutating Expert Compression:LorExperts and BTExperts

LorExperts is introduced, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant, and BTExperts is introduced, a tree organization of dominants and corrections that enables inferenc...

Inesh Chakrabarti, Sourjya Roy, B. Bao et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.