Skip to content

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Jul 2026 · arXiv.org · Vol abs/2607.05147 · 24 citations · ⚡ 8 influential · 105 references
Computer Science

TL;DR

DSpark is introduced, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification, and enables performance tiers that were previously unattainable, shifting the Pareto frontier of the DeepSeek-V4 serving system.

Abstract

Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.

View source

Similar papers

Preprint Aug 2026

CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding

This work proposes CURE, a budget-aware dynamic repair tree designed to repair errors at uncertainty focal points without incurring prohibitive tree-verification overheads, and provides a plug-and-play repair module compatible with standard parallel drafting frameworks.

Ao-Fan Liu, Jing Meng, Fangxin Liu et al. · 1 citation
Jul 2026

A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding

A unified efficiency analysis is presented showing that extending the speculation horizon can reduce rather than improve speedup when the marginal acceptance probability falls below the relative drafting cost, and SparseSpec-L, a training-free self-speculative decoding framework for long-context inference is introduced...

Yue Liu, Yuan Zeng, Min Lyu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Approximate Speculative Decoding

Approximate Speculative Decoding (ASD) is introduced, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection and reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes.

Yuan-Nuo Feng, Zegang Peng, Yu-Xin Xie et al. · 0 citations
Jul 2026

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

DFly is proposed, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel, and DFly treats verification as a shared batch-level resource.

Hong Liu, Rui Cen, Jun-Han Shi et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.