Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...
Ke Yang, Yong-Ji Gao, Xu-Shi Li et al.· 0 citations
This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.
Yan Shi, Xiao-Chao Wang, Jin-Chun Gao et al.· 1 citation
This work forms kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source, and presents a compiler-grounded and hierarchical optimization framework for Triton kernels.