Skip to content

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

Sep 2026 · 0 citations · 47 references
Computer Science

TL;DR

This work proposes ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components, which achieves speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\% over the strongest prior self-speculative baselines.

Abstract

Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule, even though the optimal draft length varies widely across requests and changes dynamically within each request. We propose ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components. First, a unified mixed forward allows drafting and verifying requests to coexist in the same batched forward pass, removing the need for global draft-verify phases. Second, a lightweight online speculation scheduler uses per-request acceptance-rate estimates and a batch-aware cost model to let each request independently choose when to verify. Third, an intra-draft refresh layer performs full attention at a single designated layer during drafting, updating the sparse context at every draft step to reduce staleness during drafting. Across three models and five reasoning and long-context benchmarks, ASPIRE achieves $1.70$-$4.58\times$ speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\%$ over the strongest prior self-speculative baselines.

View source

Similar papers

Preprint Sep 2026

NebulaSD: Many-for-Many Speculative Decoding

NebastianSD is presented, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.

Jun-Hao He, Hong-Yang Du · 0 citations
Preprint Sep 2026

Resource-Efficient Speculative Decoding for Long-Context LLM Serving

Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. Offloading KV to CPU memory relieves this pressure. However, existing offloading schemes restore the full KV history bef...

Fei Li, Song Liu, Shi-Qiang Nie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while als...

Qi-Hu Xie, Zi-Wei Li, Yi Kang · 0 citations
#artificial intelligence Preprint Sep 2026

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is...

Hao-Yuan He, Peng-Fei Liu, Si-Shi Shen et al. · 0 citations
#small language model Preprint Aug 2026

Multi-Access Speculative Inference: Uplink or Downlink?

A sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation is formulated, and a simple optimal structure is revealed that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly.

Changrui Cai, Kaibin Huang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.