Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech
Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability...