Skip to content

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

Jul 2026 · arXiv.org · Vol abs/2607.08565 · 1 citation · 80 references
Computer Science

TL;DR

SMetric addresses session-centric scheduling with differential scheduling based on two indicators derived from the request itself, the session turn and the local KV\$ hit: it schedules first-turn requests for load balance, and sticks follow-ups to the instance with the highest local hit for high KV\$ reuse.

Abstract

LLM scheduling is critical to serving, yet how well existing designs fit agentic serving--where agents, not humans, issue the requests--remains unclear. Agents shift the workload in two ways: they consume many more tokens than humans, so the cluster must provide high throughput (TPS) at low latency; and their requests reuse far more KV\$ than chat. Existing schedulers still trade off load balance against KV\$ reuse: cache-aware schedulers may crowd requests onto the few instances caching the KV\$, leaving the rest idle, while balanced schedulers may lose the opportunity for reuse, which is costly at a high reuse ratio. We thus present two key insights: (1) with a global-tier KV\$ store, pursuing load balance need not compromise KV\$ reuse, though the slower global tier must be used with care; and (2) given the agent's intra-session locality, routing requests by their sessions can balance the load with high KV\$ reuse. A key challenge in realizing session-centric scheduling is that the scheduler must identify a request's session statelessly, which is difficult for model providers serving arbitrary agents. SMetric addresses this with differential scheduling based on two indicators derived from the request itself, the session turn and the local KV\$ hit: it schedules first-turn requests for load balance, and sticks follow-ups to the instance with the highest local hit for high KV\$ reuse. As sessions differ widely in size, SMetric sticks a follow-up only if the instance can serve it within its SLO, and otherwise migrates the session to the least-loaded instance to prevent many long sessions from eventually imbalancing the load. Evaluated on real-world traces, SMetric improves the peak TPS by 9-15% under prefill-decode colocation with a provisioned global tier and the peak prefill TPS by 9% under disaggregation over state-of-the-art schedulers, also with lower latency.

View source

Similar papers

Book Open access Sep 2026

Congestion-Aware Serving of Agentic LLM Applications

Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an ela...

Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al. · 0 citations
Preprint Aug 2026

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

Cascade, an LLM serving system that estimates and continuously updates this per-request latency budget from request characteristics, KV-cache state, and current system load, and uses a single per-request budget to jointly coordinate request scheduling and KV-cache management across the memory hierarchy.

Muhammad Adnan, R. Mahapatra, Prashant J. Nair et al. · 0 citations
Preprint Sep 2026

PipeSwift: Revisiting Pipeline Parallelism for Large-Scale Completion-Oriented Agentic LLM Serving

It is shown that pipeline parallelism (PP), long overlooked because it offers little decode-latency advantage, can reduce JCT by providing a more favorable balance between prefill and decode efficiency, and PipeSwift is built, an optimized open-source pipeline-parallel runtime integrated with a tailored micro-batch par...

Shiju Wang, Fei Ren, Fang-Cheng Fu et al. · 0 citations
Preprint Aug 2026

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

A Task-Oriented Prefix-Aware Scheduler that jointly decides which agent prefixes to keep in the cache and which requests to schedule for execution and scores candidate post-decision states by trading off the expected reduction in each task's longest remaining service path against the near-term benefit of downstream pre...

Hongqiu Ni, Han Tian, Chi Zhang et al. · 2 citations
Preprint Aug 2026

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

When affinity recovers too little KV work, its residual load skew reduces or erases the improvement, so gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone is recommended.

Huang Cheng · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.