Skip to content

Generative Testing of Automated Speech Recognition Systems

Jul 2026 · arXiv.org · Vol abs/2607.09833 · 0 citations · 57 references
Computer Science

TL;DR

GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model, demonstrates that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.

Abstract

Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications. However, they remain vulnerable to adversarial manipulation, particularly in black-box settings where attacks must preserve perceptual naturalness. This work introduces GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model. Instead of perturbing waveforms directly, the approach interpolates latent representations to induce transcription errors while remaining within the manifold of natural speech. The attack is formulated as a multi-objective optimization problem balancing semantic divergence and perceptual quality. Our empirical evaluation against both white-box and black-box baselines shows that GATAS achieves a 98% success rate while producing lower distortion and higher perceptual quality, as confirmed by human studies. Despite operating without gradient access, GATAS remains competitive against white-box methods, highlighting that representation and perceptual alignment are more critical than access to model internals. Overall, our results demonstrate that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.

View source

Similar papers

Preprint Aug 2026

Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages

Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.

Varun Rai, J. PavanKumar, Sujith Pulikodan et al. · 0 citations
Open access 2026

Conditional Speaker Normalization of Vocal Tract Shape Features for Robust Synthetic Speech Detection

The rapid advancement of text-to-speech and voice conversion technologies has significantly improved the quality of synthetic speech, posing increasing challenges for developing reliable detection countermeasures. This study proposes a synthetic speech detection approach based on Log-Area Ratios (LARs) as acoustic features that provide a direct representation of vocal-tract shape. The underlying premise is that synthetic speech may exhibit inconsistencies in the speech production process that manifest as unnatural vocal-tract configurations, which can be captured through LAR features. To mitigate the speaker-dependent nature of LARs, we introduce a Conditional Speaker Normalization module that conditions the normalization process on speaker embeddings derived from either speaker verification systems or a speaker-height estimation model. These embeddings are processed through a set of fully connected layers to generate a bounded scaling vector that modulates the LAR features, improving their robustness and discriminative power in the evaluated setting. Experimental results show that the proposed approach improves performance on the FoR and Logical Access protocol of ASVspoof 2019 datasets, particularly with the speaker-height-related conditioning representation. Phoneme-level Integrated Gradients analysis indicates higher relative attribution to specific articulation categories, especially vowels. These patterns are consistent with a possible reduction in speaker-dependent articulatory variability, although the results do not establish that anatomical height information alone causes the improvement.

Hossein Fayyazi, Yasser Shekofteh · 0 citations
Preprint Aug 2026

Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models'capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

Theo Lebryk, David Ayllón, Alice Baird et al. · 0 citations
Open access Jul 2026

CAFAD: common acoustic features for adversarial audio detection

CAFAD is proposed, a plug-and-play detection framework that combines multi-domain acoustic feature fusion and temporal pyramid matching for variable-length adversarial audio detection that achieves an average detection accuracy of 99.25%, with a false positive rate of 1.00% on benign samples.

Wenjie Li, Peng-Yu Wei, Xuejing Yuan et al. · 0 citations
Preprint Aug 2026

Unsupervised Speech Recognition at the Syllable Level

Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to low-resource languages that have remained particularly difficult for prior methods. Code is publicly available\footnote{https://github.com/cactuswiththoughts/SylCipher}.

Liming Wang, Kai-Wei Chang, K. Kashino et al. · 0 citations
Conference Aug 2026

Speech-to-Plot: A Robust Voice-Driven Framework for Visual Grounding in Complex Acoustic Environments

In modern operational environments, rapid and hands-free target localization is crucial for situational awareness. However, traditional plotting systems rely on cumbersome manual interactions, and conventional multimodal algorithms degrade significantly under extreme background noise and constrained communication links. To address these challenges, we propose a novel edge-cloud collaborative Speech-to-Plot (STP) framework. The proposed system integrates a domain-adapted Automatic Speech Recognition (ASR) module—fine-tuned via a noise-injected curriculum—with a zero-shot visual grounding model to translate natural voice commands into precise spatial bounding boxes. Evaluations on a custom domain-specific dataset demonstrate that our framework exhibits graceful degradation rather than severe degradation under extreme acoustic interference, maintaining robust target semantic extraction even at 0 dB Signal-to-Noise Ratio. This reliable acoustic front-end helps prevent cascading errors in downstream cross-modal attention mechanisms, enabling accurate visual target localization. Furthermore, stress testing under simulated narrowband communication networks validates the practical engineering viability of our decoupled architecture. By offloading heavy multimodal inference to the cloud, the system mitigates computational congestion, supporting operational resilience despite the inevitable physical bandwidth limitations of field deployments.

Zhong-Hao Zhou, Hai-Lu Xin, Ping Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.