Skip to content

Author

Yan-Li Wang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the...

Kun-Ming Shao, Jie-Run Chen, Jiang-Nan Yu et al. · 0 citations
#machine learning Preprint Sep 2026

PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

At each decoding step a language model attends over the key-value (KV) cache of every earlier token, so at long context the attention call is bounded by memory bandwidth. Sparse attention reads only a subset of keys chosen by a cheap score estimate, and most methods give the unread tokens zero weight. The output then d...

Kun-Ming Shao, Jie-Run Chen, Yan-Li Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.