Book
Open access
Jul 2026
C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.
Chuheng Du, Jun-Yi Chen, Hanlin Tang et al.
· Proceedings of the 32nd ACM... · 2 citations