Skip to content
Open access

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model.

Aug 2026 · IEEE Transactions on Visualization and Computer Graphics · Vol PP, pp. 1-13 · 0 citations · 78 references
Computer Science Medicine

TL;DR

AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements and enhances usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance.

Abstract

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (CsF) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at AnyTalk.

Read PDF

Similar papers

Preprint Aug 2026

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

Wan-Animate-2 is presented, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer and achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely.

Guangyuan Wang, Liucheng Hu, Dechao Meng et al. · 2 citations
Preprint Aug 2026

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically d...

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al. · 0 citations
Sep 2026

MeshSeqGen: Zero-shot Mesh Animation based on a Large Video Generation Model.

Motion generation for 3D meshes is a fundamental task in computer animation, yet traditional methods like keyframe animation and motion capture are often costly and resource-intensive. Motivated by the progress in large-scale 2D video generation models, we introduce MeshSeqGen, a zero-shot framework that simplifies thi...

Yun-Peng Xiao, Jie Yang, Chuan Li et al. · 0 citations
Preprint Sep 2026

VISTA: Video-Injected Stylized Text-to-Animation

We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips int...

Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek · 0 citations
#artificial intelligence Preprint Sep 2026

WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...

Sangeyl Lee, Seunghyun Shin, S. Park et al. · 0 citations
Open access Sep 2026

Bridging Text and Motion: Generative AI Models for Video Synthesis

Text to video generation has advanced significantly in recent years, largely due to the development of extremely sophisticated diffusion models. In this work, we present a novel ap- proach to producing excellent video content based on descriptions by utilizing diffusion tech- niques. Using a multi-stage diffusion proce...

Mohammad Shahnawaz Shaikh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.