Skip to content

Author

Liang-Jun Feng

We have 3 of 3 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...

Ke Yang, Yong-Ji Gao, Xu-Shi Li et al. · 0 citations
Preprint Aug 2026

VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference

This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.

Yan Shi, Xiao-Chao Wang, Jin-Chun Gao et al. · 1 citation
Jul 2026

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

This work forms kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source, and presents a compiler-grounded and hierarchical optimization framework for Triton kernels.

Dongjie Chen, Ping Zhao, Bohua Zhan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.