Preprint
Aug 2026
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
Cascade, an LLM serving system that estimates and continuously updates this per-request latency budget from request characteristics, KV-cache state, and current system load, and uses a single per-request budget to jointly coordinate request scheduling and KV-cache management across the memory hierarchy.
Muhammad Adnan, R. Mahapatra, Prashant J. Nair et al.
· 0 citations