C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.