VoxTubeS is presented, a family of speaker-anonymized synthetic speech corpora designed for redistribution, comprising three method families and seven variants derived from the VoxTube corpus, which is distributed under CC BY-NC-SA 4.0.
Abstract
Large speech corpora support research, but recordings can expose speaker identity because voice remains a recognizable biometric. Meanwhile, speech data derived from media can be difficult to redistribute reliably. We present \emph{VoxTubeS}, a family of speaker-anonymized synthetic speech corpora designed for redistribution, comprising three method families and seven variants derived from the VoxTube corpus, which is distributed under CC BY-NC-SA 4.0, using 1.29M quality-filtered English utterances from 1,511 speakers. The synthesis methods span voice conversion, latent-space anonymization, and controllable text-to-speech. We evaluate VoxTubeS using utterance-level unlinkability, conversation-level linkability and singling-out, downstream speaker verification, linguistic consistency, speaker diversity, and fairness metrics for gender and accents. Our comprehensive analysis exposes a complex trade-off: stronger identity suppression often reduces linkability but sacrifices utility and population diversity, whereas speaker consistency training improves both utterance- and conversation-level privacy while retaining comparable utility and a broader speaker space. Fairness varies independently of aggregate performance. No method dominates; VoxTubeS therefore treats corpus construction as a choice among operating points that balances privacy, utility, diversity, fairness, and responsible redistribution under the source license.
Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capa...
Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak et al.· 0 citations
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary traini...
Richard Yucheng He, Bao-Dong Cao, Chen Xu et al.· 1 citation
Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 veri...
Systems for speaker anonymization obfuscate the speaker of an utterance, while maintaining its original semantic contents and prosody. Recent solutions for speaker anonymization rely on learned representations that disentangle an utterance into semantic contents and speaker properties. To anonymize an utterance, these...
Ivoline C. Ngong, Jack D'Iorio, Hailey Schoppe et al.· 0 citations
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.