Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· 0 citations· 33 references
Computer Science
TL;DR
This work presents ARK, a distributed elephant-flow path reservation mechanism that coordinates senders to choose source ports whose hashes map concurrent flows onto distinct spines, without requiring switch changes or receiver-side packet reordering.
Abstract
Disaggregated LLM inference separates the prefill and decode phases across GPU pools, generating massive KV-cache transfers. Because these transfers last hundreds of milliseconds to seconds, they behave as mega elephant flows that dominate link bandwidth utilization. In this regime, stateless ECMP can perform poorly: hash collisions may overload one spine link while leaving others idle, stretching transfer times by seconds. Yet this same persistence makes coordination practical. Since these flows are long-lived, even lightweight one-to-all coordination can be amortized over their lifetime. We present ARK, a distributed elephant-flow path reservation mechanism. ARK coordinates senders to choose source ports whose hashes map concurrent flows onto distinct spines, without requiring switch changes or receiver-side packet reordering. Packet-level RDMA simulations show that ARK reduces mean and P95 FCT by up to 27.3% and 34.0% under moderate load, and further reduces mean TTFT by up to 12.9%.
Evaluation on a 5-node GPU cluster under synthetic and production-derived churn shows that Connex reduces P99 tail spikes by up to 85% compared to NCCL-based baselines, achieves sub-second cutover, and maintains 100% goodput at moderate loads where baselines collapse to 0–28%, while incurring less than 5% steady-state...
Yanying Lin, Vincent Liu, Tao Luo et al.· Conference on Applications,...· 1 citation
DualPath is an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.
Yongtong Wu, Shaoyuan Chen, Rilin Huang et al.· Conference on Applications,...· 0 citations
Disaggregated LLM serving separates prefill and decode into distinct node pools, interposing a network fabric between the moment a key-value (KV) cache is computed and the moment it is consumed. This architectural shift invalidates a core assumption of classical cache policies: that the cost of a miss is simply recompu...
Dong Liu, Yan-Xuan Yu, Eric Jiang et al.· Proceedings of the 19th ACM...· 0 citations
Results show that object-value signals rank what to retain, while persistent responsibility determines which group bears reclamation pressure, which shows that object-value signals rank what to retain, while persistent responsibility determines which group bears reclamation pressure.
When affinity recovers too little KV work, its residual load skew reduces or erases the improvement, so gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone is recommended.
The authors' elastic KV cache lends the reserve to the KV pool during decode and returns it before prefill, driven by the scheduler's one-step-ahead view of the next batch, driven by the scheduler's one-step-ahead view of the next batch.
S. Sivashanmugam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.