The major workloads in modern large language model (LLM) serving systems have shifted from single-shot LLM calls to multi-turn conversations, where new responses are generated based on the whole conversation history across all previous turns. The hit ratio, i.e., the average fraction of KV caches accessed directly from...
He-Yuan Yao, Chu-Tong Gao, Yuan Lyu et al.· 0 citations
Modern data center servers process multiple jobs in parallel to improve performance. However, each job demands some subset of a server's resources (e.g., CPUs, memory, storage), and a set of jobs can run in parallel only if there are sufficient computational resources to meet each job's needs. Given a stream of arrivin...
This work introduces the first policy to guarantee heavy-traffic optimal mean response time in the generalized switch, the Smallest Equalizing Bucket (SEB) policy, and proves SEB's heavy-traffic optimality.
Run-Han Xie, Ziv Scully, R. Righter et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.