Aug 2026· 2 citations· ⚡ 2 influential· 27 references
Computer Science
TL;DR
ExFold is proposed, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode that calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts.
Abstract
Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts'contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughp...
Gunho Park, Kyoungho Jeun, Juntaek Oh et al.· 0 citations
ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data,...
Zu-Kang Xu, Zhi-Xiong Zhao, Xing Hu et al.· 0 citations
MetaNet is proposed, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias and provides a tunable accuracy-expert-activation trade-off on DeepSeek-MoE-16B-Chat.
Rong-Feng Wang, Shitao Weng, Zhiquan Wang et al.· 0 citations
OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs by reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit.
De-Zhi Li, Lu-Jun Li, Qi-Yuan Zhu et al.· 0 citations
LorExperts is introduced, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant, and BTExperts is introduced, a tree organization of dominants and corrections that enables inferenc...
Inesh Chakrabarti, Sourjya Roy, B. Bao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.