Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Buoy: Efficient and Effective Cache Replacement for Prefix Caching

Modern large language model (LLM) systems widely employ prefix caching to enable key-value (KV) cache reuse across different queries to minimize inference costs. At the heart of prefix caching is the replacement algorithm, which is crucial for managing limited cache space across the GPU–CPU memory hierarchy. However, t...

Liangshaowei Wang, Ran-Jun Jia, Kai Wang et al. · 0 citations
#machine learning Preprint Sep 2026

PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine.

Qiu-Yang Zhang, Kai Zhou, Kai Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.