This work proposes a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image, and develops a lightweight Face Adapter that aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training.
Abstract
Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image. A lightweight Face Adapter, together with soft-tuning of the face encoder's upper blocks, aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training. We evaluate on held-out identities from LRS3, a large-scale audiovisual corpus of English TED-talk videos. The synthesized speech is highly natural (UTMOS 3.7-4.0, matching or exceeding the 3.61 of ground truth), face-to-voice retrieval is consistently above chance, and the generated voice is consistent with the target speaker. Without any retraining, an English-trained adapter also produces fluent Spanish speech, indicating that the face-to-style mapping is largely language-agnostic.
Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.
This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.
Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu et al.· 0 citations
This work shows that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time.
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko et al.· 0 citations
NeRF-based talking head methods can render individual frames with impressive fidelity, yet the assembled videos often flicker. The reason is structural: each frame is optimized independently, so nothing in the training objective ties frame t to frame t − 1. We present TemporalTalk, which closes this gap with three trai...
Nhan T. Huynh, Nguyen N. D. Tran, Tuấn Anh Huỳnh et al.· International Conference on...· 0 citations
Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field...
Rishabh Jain, A. Papadopoulos, Zhao-Feng Lin et al.· 0 citations
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. We introduce REAL-2MIX, a Mandarin AV-TSE benchmark of jointly recorded real two-speaker mixtures with synchronized multi-view video....
Pei-Jun Yang, Zhan Jin, Xiao-Yi Qin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.