Ouroboros is proposed, a novel backdoor attack framework that leverages the ideal clean outputs of speech enhancement models as natural triggers, enabling inference-time activation without any external trigger injection.
Abstract
Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing nature of speech enhancement models. In this paper, we propose Ouroboros, a novel backdoor attack framework that leverages the ideal clean outputs of speech enhancement models as natural triggers, enabling inference-time activation without any external trigger injection. Extensive evaluations show Ouroboros achieves near-perfect attack success rates with minimal performance degradation on diverse models and datasets. Physical-world validations confirm that naturally recorded, unaltered clean audio can reliably activate the backdoor. Moreover, Ouroboros generalizes to targeted content-tampering attacks and remains effective against common filtering and finetuning defenses.
During poisoning, a trigger is injected into the forced-aligned time span of a chosen source word in the audio and replace only that word in the transcript, enabling precise semantic flips and composable sentence manipulation while avoiding many-to-one label artifacts.
Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain chall...
Yigitcan Özer, Zhe Zhang, Wan-Ying Ge et al.· 1 citation
Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...
Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, barriers encountered while building and deploying a detector are connected to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non- expert...
Anton Firc, Kamil Malinka, V. Stanek et al.· 0 citations
Deep neural network-based Voice Conversion (VC) and Text-to-Speech (TTS) models have rapidly advanced, enabling realistic voice cloning with minimal input data. Such capabilities raise serious concerns over unauthorized cloning of speaker identities and the associated privacy and security risks. Current imperceptible a...
Deepfakes have raised widespread concern owing to their threats to privacy, security, and societal trust, driving growing research interest in effective detection methods. Audio deepfake detection has moved from handcrafted-feature classifiers to end-to-end deep learning architectures and, more recently, to self-superv...
Yu-Pei Li, Yi-Xiong Fang, Jia-Hao Xue et al.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.