Skip to content

Author

Wenyuan Yu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDC

Serving offline large language model (LLM) inference workloads (e.g., log summarization and bulk translation) can consume up to 30% of GPUs in production. Despite this significant share, the characteristics of offline inference remain largely understudied. In this paper, we start by analyzing 1.5 million tasks comprisi...

Le-Ping Yang, Xue Li, Kun Qian et al. · 0 citations
#natural language process... Preprint Sep 2026

When Parallel Drafter Meets Parallel Speculative Decoding

DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative decoding (PSD) overlaps drafting with verification, yet existing methods must guess the accepted prefix and bonus token in advance: a wrong g...

Fu-Liang Liu, Xue Li, Kun Qian et al. · 0 citations
Jul 2026

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

SMetric addresses session-centric scheduling with differential scheduling based on two indicators derived from the request itself, the session turn and the local KV\$ hit: it schedules first-turn requests for load balance, and sticks follow-ups to the instance with the highest local hit for high KV\$ reuse.

Jiahao Wang, Kai-Zhan Lin, Kaixin Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.