Sep 2026· APSIPA Transactions on Signal and Information Processing· Vol 15, pp. 532-555· 0 citations· 29 references
TL;DR
This study proposes transcription quality labels (TQL), automatically derived from connectionist temporal classification scores, and confirms that synthesis quality varies systematically with TQL values at inference, demonstrating learned quality-conditioned behavior.
Abstract
Automatic speech recognition (ASR) has been increasingly adopted for generating transcriptions to train text-to-speech (TTS) models, substantially reducing manual annotation costs. However, transcription errors introduced by ASR systems inevitably degrade TTS performance, and existing approaches lack explicit quality awareness, leaving it unclear whether TTS models can learn to distinguish and respond to varying transcription quality. This study investigates whether explicit quality supervision can enable TTS models to develop quality-aware representations and achieve controllable stable synthesis. This study proposes transcription quality labels (TQL), automatically derived from connectionist temporal classification scores. During training, TQL provides explicit quality supervision, enabling the model to associate transcription quality with acoustic characteristics. During inference, setting TQL to high values guides the model toward stable synthesis that faithfully follows the input text. Under controlled conditions isolating the effect of transcription quality, our best TQL-based model achieves 8.4% word error rate compared to 15.9% for the baseline, a 47% relative reduction. While synthesis accuracy gains are significant, naturalness improvements are modest (mean opinion score of 3.66 vs. 3.55), indicating that TQL primarily enhances synthesis stability rather than perceptual quality. Further analysis confirms that synthesis quality varies systematically with TQL values at inference, demonstrating learned quality-conditioned behavior.
Synthetic speech can provide additional supervision for automatic speech recognition (ASR), but constructing useful synthetic training data requires choosing both what to synthesize and how to synthesize it. We present a phoneme-guided text-to-speech (TTS) augmentation pipeline for ASR that connects multilingual speech...
Zhen Wang, Tian-Rui Wu, Rong-Qi Han et al.· 0 citations
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while spee...
In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and further approach fully supervised performance, demonstrating that gradient-based ISL is an effective solution to expressive label scarcity in low-resource TTS.
Nicholas Sanders, G. Henter, Simon King et al.· 0 citations
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and met...
Mithilesh Vaidya, Stephen W. Bailey, Sumukh Badam et al.· 0 citations
Low-resource text-to-speech (TTS) adaptation is constrained by scarce paired data and costly manual transcription. Existing fixed-voice TTS systems can provide relatively accurate pronunciation, but their synthetic speech offers limited speaker diversity and may exhibit flat prosody. Real recordings provide natural pro...
Jia-Yi Lu, Yi-Zhong Geng, Jing-Han Yang et al.· 0 citations
Attention encoder-decoder (AED) Speech Foundation Models achieve strong ASR performance but can generate acoustically unsupported text when inputs contain no speech, weak acoustic evidence, or unreliable transcription. We propose AURA: Activation-editing with Uncertainty-Routed Adaptation, an ultra-efficient representa...
Natarajan Balaji Shankar, Zilai Wang, Zi-Han Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.