2025· Neural Information Processing Systems· pp. 151789-151811· 2 citations· 46 references
Computer Science
TL;DR
DynaPipe is a dynamic layer redistribution scheme that adaptively balances computation by predicting execution latency in real time that reduces average end-to-end request latency by 8% to 41% across diverse workloads, outperforming state-of-the-art pipeline parallelism systems.
Abstract
To accelerate large language model (LLM) inference, pipeline parallelism partitions model layers into sequential stages, each assigned to a different device for concurrent execution. However, this method often suffers from pipeline bubbles caused by imbalanced computation in the tail stage. While upstream stages focus solely on layer-forward operations, the final stage must also handle additional post-processing tasks like sampling, which introduces significant latency. This discrepancy in workload leads to pipeline misalignment, forcing upstream stages to idle and degrading overall performance. Existing frameworks typically distribute layers evenly across stages without accounting for computational load differences. To address this, we propose DynaPipe , a dynamic layer redistribution scheme that adaptively balances computation by predicting execution latency in real time. Moreover, we introduce an asynchronous key-value (KV) cache migration coordinator to enable non-blocking layer redistribution during inference. Experiments on representative LLMs demonstrate that DynaPipe reduces average end-to-end request latency by 8% to 41% across diverse workloads, outperforming state-of-the-art pipeline parallelism systems. Our implementation is publicly available at https://github.com/xhx1022/DynaPipe .
Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...
Jingqi Feng, Guang Yang, Yukai Huang et al.· Proceedings of the 32nd ACM...· 0 citations
This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.
Ying Wan, Yuchen Xu, Chuwen Zhang et al.· Conference on Applications,...· 0 citations
This study proposes Virtual Pipeline Parallelism (VPP), which keeps chunk sizes fixed and optimizes the pipeline layout through virtual stages, which improves throughput by up to 13.1% over DCPP on long sequences and 6.7% on mixed workloads, while preserving performance on short sequences.
Yan Shi, Xiao-Chao Wang, Jin-Chun Gao et al.· 1 citation
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jia-Min Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 2 citations· ⚡1
IHS-LM, a low-latency pipeline-parallel inference framework for VLM serving on heterogeneous edge devices, introduces an Intra-Batch Hybrid Scheduling method (IHS), which dynamically adjusts the prefill–decode token ratio based on token distribution and available KV cache capacity and a Bandwidth Aware Model Layer Migr...
Yun Li, Tian-Fu Pang, Zhi-Yu Cai et al.· Proceedings of the Internati...· 0 citations
DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...
Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.