Skip to content
Preprint

Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

A second, semantic draft source is proposed: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree.

Abstract

Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct drafts already in the pool, most visibly on tool-calling traffic, where a request repeats almost everything but the few values minted for it, and where one rejected token discards the correct continuation behind it. We diagnose the failure position by position across ten benchmarks and find it to be a problem of addressing rather than of coverage: on our densest tool-calling benchmark, about half of what the strongest exact-match drafter misses is present in the pool yet unreachable by exact matching. We therefore propose a second, semantic draft source: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree. In three published drafters, at matched pool and budget, it lifts accepted length by 24-29%. Oilbird reaches 4.4x autoregressive decoding speed on API-Bank, against 3.9x for the strongest training-free baseline in our harness and 2.0x for EAGLE-3.

View source

Similar papers

#natural language process... Preprint Sep 2026

DFlow: Enabling Verifier Information Flow in Block Diffusion Speculative Decoding

Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the rejected suffix, preventing the computation spent on...

Yao-Jie Zhang, Lin-Feng Zhang, Bin Cui et al. · 3 citations
Preprint Aug 2026

OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

OO-Spec is fastest among all evaluated methods in all 21 target-benchmark cells, and outperforms every evaluated released learned drafter in each comparable cell, while the same sidecar improves on ToolSpec by 34.1% on average.

Zhi-Heng Zhang, Mu-Jie Xu, Fei Sun et al. · 0 citations
#natural language process... Preprint Sep 2026

MeRoTune: RoPE-Safe Merging with a Tunable Dial

When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction...

Salman Faroz · 0 citations
Preprint Sep 2026

NebulaSD: Many-for-Many Speculative Decoding

NebastianSD is presented, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.

Jun-Hao He, Hong-Yang Du · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.