Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing...
Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...
B. Halpern, Thomas B. Tienkamp, D. Abur et al.· 0 citations
This study proposes transcription quality labels (TQL), automatically derived from connectionist temporal classification scores, and confirms that synthesis quality varies systematically with TQL values at inference, demonstrating learned quality-conditioned behavior.
Pseudo-label distillation not only transfers the performance of SSL models to a compact model but also further improves performance by leveraging available coarse labels and data augmentation.
Takuya Fujimura, Tomoki Toda· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.