Jul 2026
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
This work proposes a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image, and develops a lightweight Face Adapter that aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training.
Carlos Muñoz-Romero, Jose A. Gonzalez-Lopez
· arXiv.org · 0 citations