Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner''speech datasets are not necessarily more reliable for real-world AD detection.
This work presents the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs and reveals inherent cross-lingual correlations in prosodic structure between certain languages.
Haopeng Xie, Ismail Rasim Ulgen, Sofia Son et al.· 0 citations
DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.
Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al.· arXiv.org· 0 citations
This work revisits this design choice and proposes a sub-center modeling framework for speaker embeddings, which improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance in zero-shot voice conversion.
Ismail Rasim Ulgen, J. Hansen, Carlos Busso et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.