Buoy: Efficient and Effective Cache Replacement for Prefix Caching
Modern large language model (LLM) systems widely employ prefix caching to enable key-value (KV) cache reuse across different queries to minimize inference costs. At the heart of prefix caching is the replacement algorithm, which is crucial for managing limited cache space across the GPU–CPU memory hierarchy. However, t...