Preprint
Aug 2026
Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap
The authors' elastic KV cache lends the reserve to the KV pool during decode and returns it before prefill, driven by the scheduler's one-step-ahead view of the next batch, driven by the scheduler's one-step-ahead view of the next batch.
S. Sivashanmugam
· 0 citations