Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...
Ke Yang, Yong-Ji Gao, Xu-Shi Li et al.· 0 citations
The architecture presented in this work provides actionable insights for designing next-generation RL training systems, and introduces a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner.
Zhenyu Han, Ansheng You, Haibo Wang et al.· arXiv.org· 49 citations· ⚡8
This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.
Yan Shi, Xiao-Chao Wang, Jin-Chun Gao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.