Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
This survey examines zero-shot voice cloning through the linked views of representation, generation, control, and deployment, and identifies open problems in benchmark standardization, speaker-similarity assessment, multilingual low-resource performance, preference-aligned synthesis, and secure real-world use.
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al.· 1 citation
This work shows that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time.
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko et al.· 0 citations
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning,...
Collective-state research increasingly treats automatic transcripts as a substitute for raw audio. We show this is not free under one widely used pipeline: WhisperX omits approximately 48% of hand-labeled backchannels (“mm-hmm”, “yeah”) in close-talk multi-party recordings, and the omission replicates across held-out r...
Shan Ahmed Shaffi, M. J. Seikavandi, Anna Obara et al.· Companion Publication of the...· 1 citation
Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability...
Hounsu Kim, Joonyong Park, Yuki Saito et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.