Skip to content

Author

Qing-Sheng Yuan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Multi-Prototype Variational Information Bottleneck for SSL-Based Speech Spoofing Detection

Rapid advances in text-to-speech (TTS) and voice conversion (VC) have substantially lowered the barrier to generating zero-shot spoofed speech, posing critical challenges for identity authentication, media forensics, and speech security. Current audio deepfake detection methods typically employ large-scale self-supervised learning (SSL) models as front-end feature extractors to improve detection performance and generalizability. However, SSL representations often encode information irrelevant to spoofing discrimination, such as speaker identity, semantic content, emotional prosody, and channel conditions. These redundant factors interfere with downstream detectors’ ability to model discriminative features of spoofed speech and impair generalization to unknown attacks and cross-domain scenarios. To address this limitation, we propose the Multi-Prototype Variational Information Bottleneck (MP-VIB), a method that compresses representations to retain spoofing-relevant information while discarding spoofing-irrelevant redundancy. Critically, MP-VIB assigns a single-prototype prior for real speech to capture a compact genuine distribution. For spoofed speech, a multi-prototype mixture prior explicitly models the latent structure arising from diverse spoofing types. This asymmetric design prevents the over-compression of spoofing-related information that results from standard single-prior formulations, thereby yielding more robust latent representations for deepfake detection. Experiments with a WavLM front-end demonstrate that MP-VIB achieves a pooled equal error rate (Pooled EER) of 16.07% across 14 evaluation sets. This represents a 14.4% relative improvement in Pooled EER over the standard variational information bottleneck. On ASVspoof 2019 LA, ASVspoof 2021 LA and ASVspoof 2021 DF, the method attains EERs of 0.33%, 2.47% and 4.13%, respectively. The code is available at https://github.com/Hench-Ho/MP-VIB

Heng-Chang Hou, Ruo-Hua Zhou, Qing-Sheng Yuan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.