Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, ty...
Speech deepfake detection (SDD) models are trained on multi-attack datasets containing diverse spoofing systems, such as text-to-speech (TTS) and voice conversion (VC). In standard classifier training on multi-attack datasets, all attacks are treated as one spoofed class, and performance is reported using overall Equal...
This work presents the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs and reveals inherent cross-lingual correlations in prosodic structure between certain languages.
Haopeng Xie, Ismail Rasim Ulgen, Sofia Son et al.· 1 citation
DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.
Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al.· arXiv.org· 0 citations
This work revisits this design choice and proposes a sub-center modeling framework for speaker embeddings, which improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance in zero-shot voice conversion.
Ismail Rasim Ulgen, J. Hansen, Carlos Busso et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.