Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean are proposed, which achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.
Feng Yin, Shuai Shi, Junjie Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.