Skip to content
Preprint

RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning

Sep 2026 · 0 citations · 25 references
Computer Science

Abstract

Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.

View source

Similar papers

Review Open access 2026

Voice Cloning: A Survey of Zero-Shot and Controllable Speech Synthesis

This survey examines zero-shot voice cloning through the linked views of representation, generation, control, and deployment, and identifies open problems in benchmark standardization, speaker-similarity assessment, multilingual low-resource performance, preference-aligned synthesis, and secure real-world use.

Niraj Kumar Tiwari, Deepak Kumar, Asif Ekbal · 0 citations
#artificial intelligence Preprint Sep 2026

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al. · 1 citation
Preprint Aug 2026

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

This work shows that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time.

Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko et al. · 0 citations
#natural language process... Preprint Sep 2026

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning

TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning,...

Jian Chen, Zhang You, M. Vinton · 0 citations
Book Open access Oct 2026

Backchannel-Aware Transcription for Multiparty Conversation: ASR Omissions, Acoustic Recovery, and Cross-Corpus Transfer

Collective-state research increasingly treats automatic transcripts as a substitute for raw audio. We show this is not free under one widely used pipeline: WhisperX omits approximately 48% of hand-labeled backchannels (“mm-hmm”, “yeah”) in close-talk multi-party recordings, and the omission replicates across held-out r...

Shan Ahmed Shaffi, M. J. Seikavandi, Anna Obara et al. · 1 citation
Preprint Oct 2026

Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech

Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability...

Hounsu Kim, Joonyong Park, Yuki Saito et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.