TRAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles, and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect is introduced.
Abstract
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $\Lambda$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $\Lambda$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for reducing group word-error-rate (WER) gaps...
Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis· 0 citations
Deep neural network-based Voice Conversion (VC) and Text-to-Speech (TTS) models have rapidly advanced, enabling realistic voice cloning with minimal input data. Such capabilities raise serious concerns over unauthorized cloning of speaker identities and the associated privacy and security risks. Current imperceptible a...
At 460 h of clean LibriSpeech, the seed moves fairness metrics more than compression does on most demographic axes, and a balanced 3x3 decomposition attributes 85.3% of the variation in Fair-Speech ethnicity normalized gap to the seed against 8.3% to compression.
Srishti Ginjala, E. Fosler-Lussier, Srinivasan Parthasarathy· 0 citations
This work presents the first black-box MIA framework explicitly tailored to TTS models at both the speaker and record levels, and characterize the feasible query space and establish two criteria, scorable extent and memorization elicitation, for evaluating five representative queries.
Kun-Lin Cai, Kai-Yuan Zhang, Zihang Xiang et al.· 0 citations
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on S...
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar· 0 citations
This work empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures and proposes a two-pool allocation that normalises encoder and LLM parameters into independent pools, and shows that joint $\ell_2$ sensitivity and the original $(\varepsilon,\delta)$-DP guarantee are unchanged...
Jordi Luque, Fernando López, Aleix Sant· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.