Single-GPU long-context inference with Mixture-of-Experts (MoE) models requires spilling the key-value cache (KVCache) to CPU memory. The spilled KV serves two complementary purposes---transferring to the GPU for attention computation, or computing in-place on the CPU---which demand opposing physical states. The optimal split between them varies with workload, yet existing KVCache abstractions offer only storage semantics over a monolithic object of a single physical state, and cannot express dynamic load balancing. We propose InplaceKVCache, the first KVCache abstraction whose format fixes each byte's physical residency at write time, so that the CPU--GPU load balance can be adjusted without moving data after placement. It realizes this as a four-region layout along two dimensions---device affinity and access pattern---turning load balancing into pure scheduling. Built on this abstraction, WriteScope splits CPU--GPU shares along the sequence dimension, and a portable roofline performance model determines the optimal CPU share as sequence length evolves, with online feedback tracking CPU cost drift. On three MoE models (DeepSeek-V2-Lite, Qwen3-30B-A3B, Mixtral-8$\times$7B) with a 32~GB VRAM budget, WriteScope supports end-to-end inference at the 1M-token aggregate scale. In the long-context regime ($\ge$8K), it achieves geometric-mean speedups of $1.5\times$--$2.5\times$ on A100 and $1.4\times$--$1.7\times$ on V100 over four reproduced baselines, while vLLM, SGLang, and KTransformers fail even with a doubled KV budget. A DeepSeek-V4-Flash case study validates composition with native sparse attention.
Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency. Meanwhile, we observe that GPU utilization is high during the compute-heavy prefill phase but drops sharply during the memory-bound decode phase, where modern CPUs equipped with matrix units can achieve competitive attention throughput. We present AsymFlow, a prefill-decode disaggregated serving system that runs prefill on the GPU and decode on the CPU. AsymFlow (1) streams per-layer KV states through a shared-memory KV pool to overlap KV transfer with GPU prefill, (2) employs a task-aware online dispatcher that jointly accounts for pipeline queueing and KV capacity to prevent GPU OOM and idle bubbles, and (3) accelerates CPU decode attention with AMX. Implemented on SGLang and evaluated on a CPU-GPU platform, AsymFlow serves 32K-context workloads that can trigger OOM on a GPU-only baseline, improves throughput by 1.08–1.25 × over the GPU-only baseline, and by 1.32-2.49 × over a state-of-the-art KV-offloading method; the gains increase as GPU memory becomes tighter.
Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.