Skip to content
Preprint

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

MetaNet is proposed, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias and provides a tunable accuracy-expert-activation trade-off on DeepSeek-MoE-16B-Chat.

Abstract

Mixture-of-Experts (MoE) models route each token to a subset of expert networks, increasing capacity while keeping per-token computation sparse. In many deployed MoEs, the number of active experts is fixed across layers and tasks, although layer roles and expert redundancy vary with depth and demand varies with difficulty. Existing approaches address only part of this setting: layer-wise allocations are usually determined offline and reused for all tasks, while token-level methods vary expert activation using local routing signals without task-level context. We propose MetaNet, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias. The backbone, experts, and router remain frozen. On DeepSeek-MoE-16B-Chat, MetaNet provides a tunable accuracy-expert-activation trade-off. Relative to fixed k=6, a conservative setting activates 3.61 experts on average (40% fewer) and achieves comparable MMLU accuracy (0.489 vs. 0.474), whereas an aggressive setting activates 2.28 experts on average (62% fewer) with accuracy approximately 3.7 percentage points lower. The MMLU-trained controller also transfers to C-Eval without retraining, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.

View source

Similar papers

Preprint Aug 2026

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. We demonstrate that this uniformity is systematically suboptimal and propose MAPLE, a plug-and-play framewor...

Lie Li, Wen Li, Junxiao Shen et al. · 0 citations
Preprint Jul 2026

TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

This work introduces Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervision allocation and leverages existing routing behavior without restricting adaptation to a fixed expert subset or imposing an explicit target routing distribution.

Guanzhi Deng, Haibo Wang, Kuan Wu et al. · 0 citations
Jul 2026

Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA

Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncerta...

Tom Saliencro, Rohan Desai, Priya Nair et al. · 0 citations
Preprint Aug 2026

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compressio...

Lei Xin, Bin Gu, Peize Li et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be suff...

Dohyeon Kim, Bedionita Soro, Sung Ju Hwang · 0 citations
#machine learning Preprint Sep 2026

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

This work introduces RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private...

Duc Dm, Khai Le-Duc, D. Nguyen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.