Skip to content
Book Open access

AsymFlow: Enabling Long-Context LLM Serving via CPU-GPU Prefill-Decode Disaggregation

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 359-369 · 0 citations · 33 references

TL;DR

This work presents AsymFlow, a prefill-decode disaggregated serving system that runs prefill on the GPU and decode on the CPU, and implements on SGLang and evaluated on a CPU-GPU platform.

Abstract

Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency. Meanwhile, we observe that GPU utilization is high during the compute-heavy prefill phase but drops sharply during the memory-bound decode phase, where modern CPUs equipped with matrix units can achieve competitive attention throughput. We present AsymFlow, a prefill-decode disaggregated serving system that runs prefill on the GPU and decode on the CPU. AsymFlow (1) streams per-layer KV states through a shared-memory KV pool to overlap KV transfer with GPU prefill, (2) employs a task-aware online dispatcher that jointly accounts for pipeline queueing and KV capacity to prevent GPU OOM and idle bubbles, and (3) accelerates CPU decode attention with AMX. Implemented on SGLang and evaluated on a CPU-GPU platform, AsymFlow serves 32K-context workloads that can trigger OOM on a GPU-only baseline, improves throughput by 1.08–1.25 × over the GPU-only baseline, and by 1.32-2.49 × over a state-of-the-art KV-offloading method; the gains increase as GPU memory becomes tighter.

Read PDF

Similar papers

Preprint Sep 2026

Resource-Efficient Speculative Decoding for Long-Context LLM Serving

Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. Offloading KV to CPU memory relieves this pressure. However, existing offloading schemes restore the full KV history bef...

Fei Li, Song Liu, Shi-Qiang Nie et al. · 0 citations
#machine learning Preprint Sep 2026

PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine.

Qiu-Yang Zhang, Kai Zhou, Kai Lu et al. · 0 citations
#natural language process... Preprint Sep 2026

Deadline-Aware Adaptive Prefill Chunking for Efficient Large Language Model Serving

Continuous batching improves large language model (LLM) serving throughput, but long prompt prefills can delay decode iterations and violate inter-token latency objectives. Chunked prefill mitigates this interference, yet its chunk size is normally fixed: small chunks protect decode latency but repeatedly pay launch ov...

Si-Yu Song, Qi-Wei Bai, Jin-Bo Hao et al. · 1 citation
Book Open access Aug 2026

Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation

This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.

Ying Wan, Yuchen Xu, Chuwen Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...

Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.