Skip to content
Preprint

GhostWord: A Fine-Grained Backdoor Attack on Automatic Speech Recognition

Sep 2026 · 0 citations · 29 references
Engineering Computer Science

TL;DR

During poisoning, a trigger is injected into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts.

Abstract

Automatic Speech Recognition (ASR) systems are widely deployed in safety-critical settings but remain vulnerable to data-poisoning backdoor attacks. Existing ASR backdoors typically use phrase-level triggers paired with a fixed target sentence, creating strong artifacts (e.g., repeated transcripts or triggers placed in non-speech regions) that simple preprocessing can mitigate. We propose GhostWord, a word-level, time-localized ASR backdoor that uses codebooks mapping short ($\approx$400\,ms) acoustic triggers to target words. During poisoning, we inject a trigger into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts. Across Common Voice (v23 English, v24 Lithuanian) and multiple backbones (Whisper-Small/Medium, MMS, SpeechT5), GhostWord achieves an average attack success rate of 89.3\% and transfers across languages and models. Adapting optimization-based defenses (ABL, ANP, SAU, I-BAU) reveals a sharp robustness--accuracy trade-off: attack success drops from 89.3\% to 29.1\% while clean WER rises from 21.5\% to 45.0\%, consistent with our theoretical analysis showing that, in high-vocabulary models, backdoor suppression structurally tends to degrade clean performance. The source code is publicly available at https://github.com/rohban-lab/GhostWord

View source

Similar papers

Preprint Sep 2026

FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models

This work proposes textsc {FreqDoor], a training-time backdoor attack that implants triggers in the frequency domain for vision-language models (VLMs) and evaluates the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering.

Yasir Arafat Prodhan, Sadad Hasan, Mohammed Imamul Hassan Bhuiyan · 0 citations
Preprint Aug 2026

Ouroboros: Self-Referential Backdoor Attacks on Speech Enhancement via Clean Audio Triggers

Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing natu...

Yunkai Zhou, Yuheng Huang, Diqun Yan · 0 citations
Preprint Aug 2026

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

The perturbation-based DoS attack targeting E2E speech models is proposed, formulated as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alig...

Shuo Cheng, Kunlan Xiang, Ming-Xuan Li et al. · 0 citations
Preprint Sep 2026

GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...

Li Wang, Kun-Yu Feng, Wan Lin et al. · 0 citations
2026

Toward Harmless Ownership Verification for Automatic Speech Recognition via Synthetic Lexical Domains

Automatic speech recognition (ASR) has been widely applied in intelligent assistants and voice services, and high-performance ASR models trained on large-scale speech data have become important commercial assets. However, the protection and verification of ASR model ownership face significant challenges. Most existing...

Hanbo Cai, Peng-Cheng Zhang, Yan Xiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.