Effective request placement in Function-as-a-Service (FaaS) platforms requires timely visibility into worker state, which changes rapidly as workers create, reuse, and evict short-lived function instances. Under the conventional cloud manager–worker architecture commonly used in FaaS platforms, however, this visibility...
Seonggyu Han, Sangwoo Kim, Minho Kim et al.· Proceedings of the Internati...· 0 citations
A heterogeneous decode-phase serving system that relocates the KV cache out of GPU memory, motivated by the retrieval-based sparse attention that recent frontier LLMs adopt to serve million-token contexts and proposes KARAT, a general-purpose PNM design that is the design point meeting all four requirements.
Hyungkyu Ham, Junhyeong Bae, Seungheon Lee et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.