Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
This work proposes a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image, and develops a lightweight Face Adapter that aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training.