Aug 2026· Frontiers in Artificial Intelligence· Vol 9· 0 citations· 39 references
Medicine
TL;DR
AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network, which enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.
Abstract
Introduction The deployment of LLMs on resource-constrained hardware is hindered by the memory-intensive KV Cache mechanism. Methods We propose AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network. Results AdaK reveals estimated KV cache reductions of up to 17.9% relative to fixed-k = 2048 baselines across 16 settings on Qwen3-4B, Qwen3-8B, and Mistral-7B. Discussion AdaK's decoupled design enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.
Across a wide range of latency and peak memory constraints, MetaKV consistently outperforms the best static configuration, improving constrained success rate (CSR), the fraction of prompts answered correctly while satisfying both constraints, by approximately 0.07 on average and up to 0.135.
Michael Wang, Keith Li, Roozbeh Bostandoost· 0 citations
This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.
The challenges of the restoration-recomputation trade-off are investigated and its impact on inference performance when left unaddressed, and an I/O-aware KV-cache management policy is presented that dynamically navigates this trade-off.
Amirhossein Najafizadeh, Vasily Tarasov, Alex Merenstein et al.· Proceedings of the 18th ACM...· 0 citations
Pinch, a cache middleware that decouples clairvoyant and importance sampling at the granularity of a planning horizon, enabling planned cache retention and bounded prefetch on DNN training on large datasets.
Yu-Chen Liu, Shu Yin· Proceedings of the Internati...· 0 citations
TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...
Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al.· 0 citations
Design rules and a reproducible evaluation protocol are contributed that jointly report quality, memory, and end-to-end speed, and a foundation for automated pipeline search under realistic single-GPU constraints is provided.
Hong-Yu Yu, Yihan Shen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.