Skip to content
Conference

Fine-Tuning Text-to-Speech Models with Turkish Speech Data

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 13 references

Abstract

Developing text-to-speech (TTS) systems for a language with limited accessible speech data such as Turkish remains a challenge. This study describes a process for creating a Turkish text-to-speech system using web-scraping data to train deep learning models. The data collection approach is based on transcribing Turkish audiobook content from YouTube and converting it into a usable dataset using normalization, piece segmentation, and human annotation methods. In this study, the performances of fine-tuning KaniTTS and Dia voice models are compared with the performance of Elevenlabs voice clone. It has been observed that fine-tuned voice models with limited resources gained the ability to synthesize at the level of commercial based API voice model.

View source