Jul 2026· Signal Processing and Communications Applications Conference· pp. 1-4· 0 citations· 24 references
Abstract
When fine-tuning large language models, the assumption is that more data means better results. This principle is often extended to text-to-speech (TTS) fine-tuning, yet remains underexplored, particularly for low-resource languages where high-quality data is difficult to obtain. In this work, the "more data is better" phenomenon is investigated for Turkish TTS using XTTS v2 by incrementally increasing training data. Speech quality is evaluated using multiple metrics, including UTMOS, NISQA, and an LLM-based TTS evaluation framework. For the LLM-based evaluation, a multimodal language model (Gemini) was prompted to assess Turkish-specific pronunciation, naturalness, and synthesis artifacts on a calibrated 1-10 scale. Results suggest that for TTS fine-tuning on non-mainstream languages, modest data investments may achieve near-optimal quality.
Automatic speech recognition (ASR) systems have achieved significant progress in recent years; however, their performance remains limited for low-resource languages such as Indonesian. Multilingual ASR models are often expected to generalize across languages, yet they frequently underperform when applied to underrepresented languages without sufficient adaptation. This study presents a comparative evaluation of three ASR models—Wav2Vec 2.0, XLS-R, and XLSR-53—on Indonesian speech to analyze the impact of monolingual fine-tuning versus multilingual pretraining. The evaluation was conducted using approximately 28 hours of validated Indonesian speech from the Common Voice Corpus version 13. Model performance was assessed using word error rate (WER) without employing any external language model to ensure a fair comparison. Experimental results demonstrate that Wav2Vec 2.0, which is fine-tuned specifically for Indonesian, achieves substantially lower WER compared to the multilingual models. Qualitative analysis further confirms that multilingual models exhibit higher omission and substitution errors. These findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian. The results provide practical guidance for deploying ASR systems in low-resource language scenarios and highlight the importance of targeted model adaptation.
Juan Hebert, Amalia Zahra· Bulletin of Electrical Engin...· 0 citations
Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Qwen2-Audio large audio language model and combines them with layer-weighted WavLM SSL representations through a learned cross-attention fusion. This design captures both high-level semantic naturalness and lowlevel acoustic distortions in a unified model. On two standard benchmarks, SOMOS and BVCC, QMOS achieves competitive system-level SRCC of 0.923 and 0.921 on SOMOS and BVCC, respectively, while using no system-ID conditioning or listener embeddings. Cross-domain evaluation yields a competitive SRCC of 0.702, suggesting the learned representations generalise across acoustic domains.
Sri Ravi Sastry Kolluru, Charan Devarakonda, S. Radhe et al.· International Conference on...· 0 citations
Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.
Developing text-to-speech (TTS) systems for a language with limited accessible speech data such as Turkish remains a challenge. This study describes a process for creating a Turkish text-to-speech system using web-scraping data to train deep learning models. The data collection approach is based on transcribing Turkish audiobook content from YouTube and converting it into a usable dataset using normalization, piece segmentation, and human annotation methods. In this study, the performances of fine-tuning KaniTTS and Dia voice models are compared with the performance of Elevenlabs voice clone. It has been observed that fine-tuned voice models with limited resources gained the ability to synthesize at the level of commercial based API voice model.
Hüseyin Çakmak, Kuzey Arar, F. B. Tek· Signal Processing and Commun...· 0 citations
It is suggested that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling, in self-supervised fine-tuning.
Wangjin Zhou, Yizhou Zhang, Yichi Wang et al.· 0 citations
This paper presents a study focused on advancing Automatic Speech Recognition (ASR) for the under-resourced language Dënë Su ˛ łıné through data-centric approaches. We explore multiple strategies to enhance the quality of training data—both audio recordings and tran-scriptions—to address the challenges posed by mixed-quality datasets. Our experiments investigate which data preparation techniques most effectively improve ASR performance in this context. Our findings show that reducing spelling variants of the same lexeme in the corpus significantly improves model generalization, resulting in a substantial increase in recognition accuracy. Additionally, we demonstrate that increasing manually reviewed transcriptions consistently improves word and character error rates, while audio enhancement slightly reduces performance, highlighting the complex trade-offs in low-resource ASR development.
Olga Kriukova, O. Lovick, Antti Arppe· Proceedings of the Sixth Wor...· 0 citations