Skip to content
Preprint

VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference

Aug 2026 · 1 citation · 34 references
Computer Science

TL;DR

This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.

Abstract

Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher attention costs, leading to pipeline bubbles. Existing approaches mitigate this imbalance through dynamic chunk resizing (Dynamic CPP, DCPP), but our measurements show that this trades scheduling overhead for load balancing, which becomes unfavorable on long sequences. In this study, we propose Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages. A V-shaped virtual-stage traversal overlaps each chunk's expensive middle stages with the lighter head and tail stages of its neighbors, while asynchronous communication and pipelined packing further reduce communication stalls and cross-request drain bubbles. We implement VPP in vLLM-Ascend and evaluate it on three MoE-based LLMs with sequences up to 1M tokens on 16 Ascend 910C NPUs. VPP improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences. On a 512K-token DeepSeek-V3.1 prefill workload, VPP reduces the pipeline bubble ratio from 6.4% to 0.1%, achieving a 98.0% reduction compared with DCPP.

View source

Similar papers

#machine learning Preprint Sep 2026

Spexis: Speculative Lookahead Scheduling for LLM Inference

Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism, and uses lookahead scheduling to predict speculation quality and future memory pressure, to reduce wasted speculation, KV-cache eviction, and recomputation.

Hyungyu Jung, Jaehyeok Yu, Hoon Choi et al. · 0 citations
Book Open access Sep 2026

AsymFlow: Enabling Long-Context LLM Serving via CPU-GPU Prefill-Decode Disaggregation

This work presents AsymFlow, a prefill-decode disaggregated serving system that runs prefill on the GPU and decode on the CPU, and implements on SGLang and evaluated on a CPU-GPU platform.

Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al. · 0 citations
Preprint Sep 2026

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap

Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...

Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al. · 0 citations
Book Open access Aug 2026

OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving

Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...

Jingqi Feng, Guang Yang, Yukai Huang et al. · 0 citations
#natural language process... Preprint Sep 2026

Deadline-Aware Adaptive Prefill Chunking for Efficient Large Language Model Serving

Continuous batching improves large language model (LLM) serving throughput, but long prompt prefills can delay decode iterations and violate inter-token latency objectives. Chunked prefill mitigates this interference, yet its chunk size is normally fixed: small chunks protect decode latency but repeatedly pay launch ov...

Si-Yu Song, Qi-Wei Bai, Jin-Bo Hao et al. · 1 citation
Preprint Sep 2026

Resource-Efficient Speculative Decoding for Long-Context LLM Serving

Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. Offloading KV to CPU memory relieves this pressure. However, existing offloading schemes restore the full KV history bef...

Fei Li, Song Liu, Shi-Qiang Nie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.