Skip to content

Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation

Sep 2026 · 0 citations · 16 references
Computer Science

TL;DR

The results show that low-rank projection adaptation can recover held-out quality across fixed cache formats, while perplexity recovery need not restore long-context retrieval.

Abstract

Low-bit key--value (KV) caches reduce the memory required for autoregressive decoding, but the resulting quality loss depends on the model and quantizer. We keep the quantizer fixed and distill the floating-cache model's behavior into low-rank Q/K/V projection updates while the student executes a physically packed incremental cache. Across three seeds, 4-bit affine-cache adapters recover $54.24\%\pm2.47\%$ of the held-out perplexity gap on TinyLlama-1.1B and $75.96\%\pm4.04\%$ on Gemma-4-12B. On the same frozen NF4 Llama-3.1-8B base, one validation-selected run per quantizer recovers $60.42\%$ under KIVI K2V2 and $37.61\%$ under KVarN K4V2, while preserving 180-case associative retrieval. Gemma's score on an official 4K/8K RULER subset rises from 42.80 with the unadapted 4-bit cache to 48.33 after adaptation (46.15 floating), with substantial task heterogeneity. Finally, a 2-bit rank--token sweep reduces TinyLlama's 2-bit PPL from 576.10 to $11.4000\pm0.0059$ across three seeds, versus 10.3988 floating, but restores only 11--12 of 180 retrieval cases. These results show that low-rank projection adaptation can recover held-out quality across fixed cache formats, while perplexity recovery need not restore long-context retrieval.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Output-Aware Rotation for INT2 KV-Cache Quantization

OptR is proposed, an output-aware rotation method that minimizes post-W_O$ attention-output error and applies an attention-equivalent key reparameterization to reduce large channel-wise offsets without changing the softmax distribution.

Vincent-Daniel Yun, Woo-Sang Lim, Minsoo Cheong et al. · 2 citations · ⚡1
Preprint Aug 2026

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

This paper derives closed-form optimal transforms for keys and values from calibration statistics, under a high-resolution model, and shows that the optimal key transform is not orthogonal and satisfies a generalized Parseval relation: the attention-aware distortion becomes mean-squared error in the transform domain.

Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez et al. · 2 citations
#machine learning Preprint Aug 2026

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

SemKV preserves every token, ranks tokens by a model-internal score, and assigns two adjacent above-cliff precisions, achieving a measured 6.0x storage reduction with no statistically detectable quality difference from full KV (n=900, three seeds), and outperforming FP16 token pruning granted a 1.5x larger memory budge...

D. Lee, Do-Hyung Kim, Jae-Hong Kim · 0 citations
Preprint Aug 2026

KV Cache Compression Through the Lens of Transform Coding

The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the...

Hannah Laus, C. M. Verdun, Hao Wang et al. · 1 citation
Preprint Aug 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.

Joao V. Cavalcanti, Ashia C. Wilson · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.