Skip to content

DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting

Jul 2026 · arXiv.org · Vol abs/2607.07409 · 2 citations · 45 references
Computer Science

TL;DR

DeLS-Spec is proposed, a decoupled long-short context speculative decoding method that treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert and consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.

Abstract

Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash further improve drafting efficiency by predicting an entire block in one pass, but their position-wise predictions lack explicit intra-block causal conditioning. Recent methods such as Domino and DSpark attempt to introduce such causality into block-parallel drafting, but they require training the draft model from scratch, which limits their flexibility and increases training cost. We propose DeLS-Spec, a decoupled long-short context speculative decoding method. DeLS-Spec treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert. The local head can be trained independently with a standard next-token prediction objective, without joint training with the target model or the DFlash backbone, leading to extremely low training cost. At inference time, DeLS-Spec combines long-context and short-context logits, and the local head is not tied to a specific DFlash checkpoint, making the method more modular and flexible. Experiments on Qwen3 models show that DeLS-Spec consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.

View source

Similar papers

Preprint Aug 2026

DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding

This work proposes a dependent block drafter based on a low-rank latent mixture over token positions, complemented by an acceptance-oriented training objective that directly targets the expected verified length.

Amirmohammad Karimi, Chao Gao, Negar Hassanpour · 1 citation
#artificial intelligence Preprint Sep 2026

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is...

Hao-Yuan He, Peng-Fei Liu, Si-Shi Shen et al. · 0 citations
#natural language process... Preprint Aug 2026

LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

LiLiCorr, a Lightweight Likelihood-based model that Correlates the per-position marginal distributions a drafter already produces, is introduced, a Lightweight Likelihood-based model that Correlates the per-position marginal distributions a drafter already produces.

M. Rusanovsky, Yoav Miron, Roy Uziel et al. · 2 citations
#natural language process... Preprint Aug 2026

TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding

Speculative decoding accelerates language-model inference by letting a cheap drafter propose tokens that the target model verifies in parallel. Recent block drafters make drafting nearly free: a single backbone pass emits an entire block of draft tokens. Draft trees promise a further gain -- several alternative continu...

Hua-Peng Zhou, Hua-Yu Wang, Xin-Yu Wang · 1 citation
#natural language process... Preprint Sep 2026

DEdit: Iterative Draft Editing for Speculative Decoding

Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes...

Long-Xuan Yu, Bingsen Chen, Peng Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor i...

Hao-Hui Zhang, Ke-Yu Chen, Hao-Cheng Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.