Skip to content

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

A router is studied that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class to match the goodput of round robin using six GPUs instead of seven.

Abstract

Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this separation fast, but routing still determines which instances handle each request. We study a router that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class. We develop the policy in a discrete event simulator and validate it on eight NVIDIA A40 GPUs, each running a vLLM engine, with NIXL transferring KV caches between pools. All workloads run at measured saturation. Across three mixed, bursty arrival traces, the calibrated router achieves the highest mean goodput at 0.864, compared with 0.835 to 0.847 for round robin, least loaded, and a length heuristic. It also shows the lowest variance across traces. It beats round robin and the length heuristic on all three traces and least loaded on two. On the third, it trails by 0.003, within run to run noise. Hardware calibration matters: simulator derived constants cost 4.5 goodput points and roughly 40 percent of the tail latency advantage, reducing the scorer to little more than queue counting. Benefits grow with decode pool size and traffic heterogeneity but disappear in pools with three instances, where queue counts are often enough. Under extreme scarcity, greedy cost minimization concentrates requests on the cheapest scored instance, and blind spreading performs better. With calibrated costs, the learned router matches the goodput of round robin using six GPUs instead of seven.

View source

Similar papers

Preprint Aug 2026

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

When affinity recovers too little KV work, its residual load skew reduces or erases the improvement, so gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone is recommended.

Huang Cheng · 1 citation
Book Open access Aug 2026

ARK: Avoiding Routing Collisions for KV Cache Transfer in Disaggregated LLM Inference

This work presents ARK, a distributed elephant-flow path reservation mechanism that coordinates senders to choose source ports whose hashes map concurrent flows onto distinct spines, without requiring switch changes or receiver-side packet reordering.

Hung-Chun Lin, Ting-Wei Hsu, Chung-En Ho et al. · 0 citations
Preprint Aug 2026

MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory Architecture

A ReRAM near-memory architecture that keeps expert weights resident behind high-bandwidth local reads and recovers occupancy with bounded core-local multicast pooling, coactivation-aware placement, and load-aware fetch, and sizes each communication level from induced demand is presented.

Kun-Ming Shao, Ming Zeng, Xin Yuan et al. · 0 citations
Preprint Sep 2026

MoSE: Mode-Switching Expander for Mixed LLM Training and Inference

AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static spa...

Fan Yang, Ying Zhou, Bing-Lei Wang et al. · 0 citations
Book Open access Sep 2026

To Keep or Not to Keep: Learning KV Cache Retention in Disaggregated LLM Serving Systems

Disaggregated LLM serving separates prefill and decode into distinct node pools, interposing a network fabric between the moment a key-value (KV) cache is computed and the moment it is consumed. This architectural shift invalidates a core assumption of classical cache policies: that the cost of a miss is simply recompu...

Dong Liu, Yan-Xuan Yu, Eric Jiang et al. · 0 citations
#machine learning Preprint Sep 2026

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent optimization. We measure what it costs in reproducibility, and find that the cost rises sharply with weight quantizatio...

Aditi Patodiya · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.