A second, semantic draft source is proposed: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree.
Abstract
Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct drafts already in the pool, most visibly on tool-calling traffic, where a request repeats almost everything but the few values minted for it, and where one rejected token discards the correct continuation behind it. We diagnose the failure position by position across ten benchmarks and find it to be a problem of addressing rather than of coverage: on our densest tool-calling benchmark, about half of what the strongest exact-match drafter misses is present in the pool yet unreachable by exact matching. We therefore propose a second, semantic draft source: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree. In three published drafters, at matched pool and budget, it lifts accepted length by 24-29%. Oilbird reaches 4.4x autoregressive decoding speed on API-Bank, against 3.9x for the strongest training-free baseline in our harness and 2.0x for EAGLE-3.
Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the rejected suffix, preventing the computation spent on...
Yao-Jie Zhang, Lin-Feng Zhang, Bin Cui et al.· 3 citations
This analysis shows that verifier skipping is a useful new lossy axis and, surprisingly, its key challenge is prefix scheduling rather than token prediction alone.
Carryover Drafting is introduced, a parallel draft--verify--draft training that exposes the drafter to inference-aligned rejected states while preserving parallelism across training positions.
Jahyun Koo, Sunghyeon Woo, J. Kil et al.· 1 citation
OO-Spec is fastest among all evaluated methods in all 21 target-benchmark cells, and outperforms every evaluated released learned drafter in each comparable cell, while the same sidecar improves on ToolSpec by 34.1% on average.
Zhi-Heng Zhang, Mu-Jie Xu, Fei Sun et al.· 0 citations
When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction...
NebastianSD is presented, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.
Jun-Hao He, Hong-Yang Du· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.