Skip to content
Preprint

VoxTubeS: Distributable Speaker-Anonymized Synthetic Speech Corpora and Their Analysis

Sep 2026 · 0 citations · 29 references
Engineering

TL;DR

VoxTubeS is presented, a family of speaker-anonymized synthetic speech corpora designed for redistribution, comprising three method families and seven variants derived from the VoxTube corpus, which is distributed under CC BY-NC-SA 4.0.

Abstract

Large speech corpora support research, but recordings can expose speaker identity because voice remains a recognizable biometric. Meanwhile, speech data derived from media can be difficult to redistribute reliably. We present \emph{VoxTubeS}, a family of speaker-anonymized synthetic speech corpora designed for redistribution, comprising three method families and seven variants derived from the VoxTube corpus, which is distributed under CC BY-NC-SA 4.0, using 1.29M quality-filtered English utterances from 1,511 speakers. The synthesis methods span voice conversion, latent-space anonymization, and controllable text-to-speech. We evaluate VoxTubeS using utterance-level unlinkability, conversation-level linkability and singling-out, downstream speaker verification, linguistic consistency, speaker diversity, and fairness metrics for gender and accents. Our comprehensive analysis exposes a complex trade-off: stronger identity suppression often reduces linkability but sacrifices utility and population diversity, whereas speaker consistency training improves both utterance- and conversation-level privacy while retaining comparable utility and a broader speaker space. Fairness varies independently of aggregate performance. No method dominates; VoxTubeS therefore treats corpus construction as a choice among operating points that balances privacy, utility, diversity, fairness, and responsible redistribution under the source license.

View source

Similar papers

Preprint Aug 2026

Your Voice Cloning System is Secretly a Voice Anonymizer

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capa...

Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak et al. · 0 citations
#natural language process... Preprint Sep 2026

ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary traini...

Richard Yucheng He, Bao-Dong Cao, Chen Xu et al. · 1 citation
#natural language process... Preprint Sep 2026

VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching

Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 veri...

M. Hoàng, Thai Le · 0 citations
Preprint Aug 2026

DP-VOXLET: Provable Speaker Anonymization for Disentangled Speech Representations

Systems for speaker anonymization obfuscate the speaker of an utterance, while maintaining its original semantic contents and prosody. Recent solutions for speaker anonymization rely on learned representations that disentangle an utterance into semantic contents and speaker properties. To anonymize an utterance, these...

Ivoline C. Ngong, Jack D'Iorio, Hailey Schoppe et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.