Skip to content
Open access

AdaK: adaptive KV cache budget estimation framework for analyzing long-context large language model inference

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 39 references
Medicine

TL;DR

AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network, which enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.

Abstract

Introduction The deployment of LLMs on resource-constrained hardware is hindered by the memory-intensive KV Cache mechanism. Methods We propose AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network. Results AdaK reveals estimated KV cache reductions of up to 17.9% relative to fixed-k = 2048 baselines across 16 settings on Qwen3-4B, Qwen3-8B, and Mistral-7B. Discussion AdaK's decoupled design enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

Across a wide range of latency and peak memory constraints, MetaKV consistently outperforms the best static configuration, improving constrained success rate (CSR), the fraction of prompts answered correctly while satisfying both constraints, by approximately 0.07 on average and up to 0.135.

Michael Wang, Keith Li, Roozbeh Bostandoost · 0 citations
Preprint Aug 2026

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Asaad Althoubi · 0 citations
Book Open access Sep 2026

LLM KV-cache: To Restore or To Recompute, That Is the Question

The challenges of the restoration-recomputation trade-off are investigated and its impact on inference performance when left unaddressed, and an I/O-aware KV-cache management policy is presented that dynamically navigates this trade-off.

Amirhossein Najafizadeh, Vasily Tarasov, Alex Merenstein et al. · 0 citations
Book Open access Sep 2026

PINCH: Predictive Importance-Sampling-Informed Cache for I/O-Bound DNN Training

Pinch, a cache middleware that decouples clairvoyant and importance sampling at the granularity of a planning horizon, enabling planned cache retention and bounded prefetch on DNN training on large datasets.

Yu-Chen Liu, Shu Yin · 0 citations
#machine learning Preprint Sep 2026

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...

Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.