Skip to content

Author

Kui Luo

We have 3 of 5 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...

Ke Yang, Yong-Ji Gao, Xu-Shi Li et al. · 0 citations

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

The architecture presented in this work provides actionable insights for designing next-generation RL training systems, and introduces a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner.

Zhenyu Han, Ansheng You, Haibo Wang et al. · 49 citations · ⚡8
Preprint Aug 2026

VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference

This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.

Yan Shi, Xiao-Chao Wang, Jin-Chun Gao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.