KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching t...
Minsoo Cheong, Woo-Sang Lim, Vincent-Daniel Yun et al.· 0 citations
Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
Vincent-Daniel Yun, Woo-Sang Lim, Haneul Yoo et al.· 0 citations
The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though...
Vincent-Daniel Yun, Woo-Sang Lim, Minsoo Cheong et al.· 1 citation
It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.
Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al.· arXiv.org· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.