Skip to content

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

Jul 2026 · arXiv.org · Vol abs/2607.26742 · 0 citations · 31 references
Computer Science Engineering

TL;DR

This work proposes a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image, and develops a lightweight Face Adapter that aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training.

Abstract

Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image. A lightweight Face Adapter, together with soft-tuning of the face encoder's upper blocks, aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training. We evaluate on held-out identities from LRS3, a large-scale audiovisual corpus of English TED-talk videos. The synthesized speech is highly natural (UTMOS 3.7-4.0, matching or exceeding the 3.61 of ground truth), face-to-voice retrieval is consistently above chance, and the generated voice is consistent with the target speaker. Without any retraining, an English-trained adapter also produces fluent Spanish speech, indicating that the face-to-style mapping is largely language-agnostic.

View source

Similar papers

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations
#small language model Preprint Sep 2026

Towards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval

This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.

Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu et al. · 0 citations
Preprint Aug 2026

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

This work shows that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time.

Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko et al. · 0 citations
Conference Aug 2026

Integrating Temporal Supervision and Self-Attention for Audio-Driven Head Synthesis

NeRF-based talking head methods can render individual frames with impressive fidelity, yet the assembled videos often flicker. The reason is structural: each frame is optimized independently, so nothing in the training objective ties frame t to frame t − 1. We present TemporalTalk, which closes this gap with three trai...

Nhan T. Huynh, Nguyen N. D. Tran, Tuấn Anh Huỳnh et al. · 0 citations
Preprint Sep 2026

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field...

Rishabh Jain, A. Papadopoulos, Zhao-Feng Lin et al. · 0 citations
Preprint Aug 2026

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. We introduce REAL-2MIX, a Mandarin AV-TSE benchmark of jointly recorded real two-speaker mixtures with synchronized multi-view video....

Pei-Jun Yang, Zhan Jin, Xiao-Yi Qin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.