Skip to content

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Sep 2026 · 1 citation · 62 references
Computer Science

TL;DR

This work develops a novel multi-draft speculative sampling algorithm based on Poisson processes that maintains both watermark strength and sampling efficiency, and is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency.

Abstract

Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.

View source

Similar papers

Preprint Sep 2026

ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation

ROSETTA is proposed, a hybrid CKKS/TFHE framework that overcomes inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage and achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.

Jiang-Rui Yu, Bao-Sheng Zhang, Liang Kong et al. · 0 citations
#machine learning Preprint Sep 2026

Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference

We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token agains...

Simone Ceppi, Ignacio Sanchez · 0 citations
#artificial intelligence Preprint Aug 2026

OpenStamp: A Watermark for Open-Source Language Models

This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior m...

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi · 0 citations
#machine learning Preprint Sep 2026

AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks

With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of h...

Thibaud Gloaguen, Robin Staab, Martin T. Vechev · 0 citations
#artificial intelligence Preprint Sep 2026

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is...

Hao-Yuan He, Peng-Fei Liu, Si-Shi Shen et al. · 0 citations
#machine learning Preprint Aug 2026

Minimax bounds for watermarked and masked recursive discrete distribution estimation

This work provides a lower bound that shows that it is impossible to improve performance by adding watermarks unless the false negative rate of detection also vanishes, and shows that in most regimes, the worst-case losses of a sequence of simple deterministic estimators match the corresponding lower bounds up to const...

Millen Kanabar, Michael Gastpar · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.