Skip to content

When Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models

Sep 2026 · 0 citations · 68 references
Computer Science

TL;DR

TRAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles, and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect is introduced.

Abstract

Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $\Lambda$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $\Lambda$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.

View source

Similar papers

#natural language process... Preprint Sep 2026

A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models

Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for reducing group word-error-rate (WER) gaps...

Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis · 0 citations
Preprint Sep 2026

Silence-the-Mimic: Accelerating Imperceptible Perturbation Generation Against Voice Cloning

Deep neural network-based Voice Conversion (VC) and Text-to-Speech (TTS) models have rapidly advanced, enabling realistic voice cloning with minimal input data. Such capabilities raise serious concerns over unauthorized cloning of speaker identities and the associated privacy and security risks. Current imperceptible a...

Run-Qiu Xu · 0 citations
#natural language process... Preprint Sep 2026

Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation

At 460 h of clean LibriSpeech, the seed moves fairness metrics more than compression does on most demographic axes, and a balanced 3x3 decomposition attributes 85.3% of the variation in Fair-Speech ethnicity normalized gap to the seed against 8.3% to compression.

Srishti Ginjala, E. Fosler-Lussier, Srinivasan Parthasarathy · 0 citations
#machine learning Preprint Sep 2026

Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

This work presents the first black-box MIA framework explicitly tailored to TTS models at both the speaker and record levels, and characterize the feasible query space and establish two criteria, scorable extent and memorization elicitation, for evaluating five representative queries.

Kun-Lin Cai, Kai-Yuan Zhang, Zihang Xiang et al. · 0 citations
#natural language process... Preprint Sep 2026

Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs

Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on S...

Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar · 0 citations
#natural language process... Preprint Sep 2026

Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

This work empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures and proposes a two-pool allocation that normalises encoder and LLM parameters into independent pools, and shows that joint $\ell_2$ sensitivity and the original $(\varepsilon,\delta)$-DP guarantee are unchanged...

Jordi Luque, Fernando López, Aleix Sant · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.