Skip to content
Preprint

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

The perturbation-based DoS attack targeting E2E speech models is proposed, formulated as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss.

Abstract

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs. Moreover, prior studies on speech model security mainly focus on ASR or TTS systems, leaving the DoS vulnerability of E2E speech LLMs largely unexplored. To address this gap, we propose the perturbation-based DoS attack targeting E2E speech models. Instead of inducing long outputs through prompt manipulation, our method optimizes imperceptible acoustic perturbations to directly influence the model's autoregressive generation process while preserving the original input length. Specifically, we formulate the attack as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss. To further improve stealthiness, we employ voice activity detection (VAD) to inject perturbations only into voiced regions. Extensive experiments on three open-source E2E speech LLMs demonstrate that our method achieves stable attack success rate while significantly increasing generation length and GPU resource consumption, revealing security risks in modern ALLMs.

View source

Similar papers

Preprint Sep 2026

GhostWord: A Fine-Grained Backdoor Attack on Automatic Speech Recognition

During poisoning, a trigger is injected into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts.

Mojtaba Nafez, Mobina Poulaei, Kiarash Kiani Feriz et al. · 0 citations
Preprint Sep 2026

Not All Attacks Are Learned Equally in Speech Deepfake Detection

A replay-regularized, attack-aware curriculum that steps exposure based on measured attack influence is proposed that shows improved overall robustness and reduced attack-level imbalance compared with standard multi-attack training.

Avantika Singh, Aurosweta Mahapatra, Ismail Rasim Ulgen et al. · 0 citations
Preprint Aug 2026

Robust Context-Aware Detection of Malicious Instructions in Text

The proposed approach for malicious sentence classification that is both context- and query-aware and outperforms state-of-the-art IPI defense baselines under static attacks, while in the case of adaptive attacks, the AT variants provide significantly higher utility, lower attack success rate, and often both.

Buzhao Liu, Xinhang Ma, Yevgeniy Vorobeychik · 1 citation
Preprint Sep 2026

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as content moderation and video analytics. To process long videos efficiently, VideoLLMs rely on frame sampling, token compression, and modality fusion, which together form an observation pipeline that reduces the raw...

Bang-Shuo Zhu, Wei Song, Yu-Xin Cao et al. · 0 citations
2026

Random Character-Level Perturbations Amplify LLM Jailbreak Attacks

This work finds that models cannot reliably reconstruct the original meaning and layer-wise probe classifiers fail to detect the harmful intent of perturbed prompts, and perturbations can occasionally reduce attack success by inducing off-topic or incoherent responses.

Shuyi Yu, Zhe Cao, Kohei Tsuji et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.