During poisoning, a trigger is injected into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts.
Abstract
Automatic Speech Recognition (ASR) systems are widely deployed in safety-critical settings but remain vulnerable to data-poisoning backdoor attacks. Existing ASR backdoors typically use phrase-level triggers paired with a fixed target sentence, creating strong artifacts (e.g., repeated transcripts or triggers placed in non-speech regions) that simple preprocessing can mitigate. We propose GhostWord, a word-level, time-localized ASR backdoor that uses codebooks mapping short ($\approx$400\,ms) acoustic triggers to target words. During poisoning, we inject a trigger into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts. Across Common Voice (v23 English, v24 Lithuanian) and multiple backbones (Whisper-Small/Medium, MMS, SpeechT5), GhostWord achieves an average attack success rate of 89.3\% and transfers across languages and models. Adapting optimization-based defenses (ABL, ANP, SAU, I-BAU) reveals a sharp robustness--accuracy trade-off: attack success drops from 89.3\% to 29.1\% while clean WER rises from 21.5\% to 45.0\%, consistent with our theoretical analysis showing that, in high-vocabulary models, backdoor suppression structurally tends to degrade clean performance. The source code is publicly available at https://github.com/rohban-lab/GhostWord
This work proposes textsc {FreqDoor], a training-time backdoor attack that implants triggers in the frequency domain for vision-language models (VLMs) and evaluates the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering.
Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing natu...
The perturbation-based DoS attack targeting E2E speech models is proposed, formulated as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alig...
Shuo Cheng, Kunlan Xiang, Ming-Xuan Li et al.· 0 citations
Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...
Automatic speech recognition (ASR) has been widely applied in intelligent assistants and voice services, and high-performance ASR models trained on large-scale speech data have become important commercial assets. However, the protection and verification of ASR model ownership face significant challenges. Most existing...
Hanbo Cai, Peng-Cheng Zhang, Yan Xiao et al.· IEEE Transactions on Audio,...· 0 citations
Across systems, WER has little rank agreement with CTEM (Spearman $\rho=-0.28$) or TSR (Spearman $\rho=-0.28$), and even the strongest system leaves nearly one-third of recordings with an unrecovered critical value.
Tyler Baumgartner, Brandon Tai, Lisa Kaelin-Martin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.