Skip to content

Similar papers

Open access Aug 2026

AdaK: adaptive KV cache budget estimation framework for analyzing long-context large language model inference

AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network, which enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.

Tian-Jun Shao · 0 citations
#machine learning Preprint Sep 2026

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

Across a wide range of latency and peak memory constraints, MetaKV consistently outperforms the best static configuration, improving constrained success rate (CSR), the fraction of prompts answered correctly while satisfying both constraints, by approximately 0.07 on average and up to 0.135.

Michael Wang, Keith Li, Roozbeh Bostandoost · 0 citations
#artificial intelligence Preprint Sep 2026

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

This work introduces HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths and proposes SeqCalib as the core policy-generation algorithm in HeadWiseKV.

Ren-Jie Xie, Jun-Cheng Yang, Ao-Ting Hu et al. · 0 citations
#machine learning Preprint Sep 2026

PatchKV: Weight-Space Compensation of KV Cache

Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework...

Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi et al. · 0 citations
Book Open access Sep 2026

Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference

This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.

Zong-Xing Zhao, Xia-Qing Li, Ze-Kai Meng et al. · 0 citations
Book Open access Aug 2026

DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM Serving

DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...

Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.