Skip to content

Controlled Decoding Attacks on Black-Box LLMs

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.

Abstract

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.

View source

Similar papers

Preprint Sep 2026

FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation

As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting...

Zi-Han Wang, Rui Zhang, Xin-Yuan Qian et al. · 0 citations
#natural language process... Preprint Sep 2026

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is introduced, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost and doubles as a free, always-on signal for detecting backdoored LLM behavior.

Wen Rui, Ahmed Salem, Andrew Paverd et al. · 0 citations
Preprint Aug 2026

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproportionately correlated with certain tokens (e.g., end...

Hua-Kang Lin, Tian-Cheng Zheng, Ming-Xuan Sun et al. · 0 citations
Preprint Sep 2026

Practical Secrets Extraction against Black-box LLMs

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extracti...

Shi-Qian Zhao, Si-Wei Jiang, Xin-Feng Li et al. · 0 citations
Preprint Sep 2026

Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams

Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An of...

Ming-Yuan Li, Yan-Na Jiang, Guang-Sheng Yu et al. · 0 citations
#natural language process... Preprint Aug 2026

Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

A validity-gated result contract is introduced that distinguishes an observed dissimilarity from an uninformative measurement caused by missing usage data, rate limits, or endpoint policy, and validates token-count consistency as a fingerprint of a shared tokenization stack, but rejects its use as a standalone necessar...

Bolin Chen · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.