Skip to content

Author

Ismail Rasim Ulgen

We have 5 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, ty...

Aurosweta Mahapatra, Xiutian Zhao, Shreeram Suresh Chandra et al. · 1 citation
Preprint Sep 2026

Not All Attacks Are Learned Equally in Speech Deepfake Detection

Speech deepfake detection (SDD) models are trained on multi-attack datasets containing diverse spoofing systems, such as text-to-speech (TTS) and voice conversion (VC). In standard classifier training on multi-attack datasets, all attacks are treated as one spoofed class, and performance is reported using overall Equal...

Avantika Singh, Aurosweta Mahapatra, Ismail Rasim Ulgen et al. · 0 citations
#natural language process... Preprint Aug 2026

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

This work presents the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs and reveals inherent cross-lingual correlations in prosodic structure between certain languages.

Haopeng Xie, Ismail Rasim Ulgen, Sofia Son et al. · 1 citation

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.

Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al. · 0 citations
#machine learning Preprint Jul 2024

Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity

This work revisits this design choice and proposes a sub-center modeling framework for speaker embeddings, which improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance in zero-shot voice conversion.

Ismail Rasim Ulgen, J. Hansen, Carlos Busso et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.