Aug 2026· IEEE Transactions on Visualization and Computer Graphics· Vol PP, pp. 1-13· 0 citations· 78 references
Computer ScienceMedicine
TL;DR
AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements and enhances usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance.
Abstract
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (CsF) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at AnyTalk.
Wan-Animate-2 is presented, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer and achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely.
Guangyuan Wang, Liucheng Hu, Dechao Meng et al.· 2 citations
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically d...
Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al.· 0 citations
Motion generation for 3D meshes is a fundamental task in computer animation, yet traditional methods like keyframe animation and motion capture are often costly and resource-intensive. Motivated by the progress in large-scale 2D video generation models, we introduce MeshSeqGen, a zero-shot framework that simplifies thi...
Yun-Peng Xiao, Jie Yang, Chuan Li et al.· IEEE Transactions on Visuali...· 0 citations
We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips int...
Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek· 0 citations
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...
Sangeyl Lee, Seunghyun Shin, S. Park et al.· 0 citations
Text to video generation has advanced significantly in recent years, largely due to the development of extremely sophisticated diffusion models. In this work, we present a novel ap- proach to producing excellent video content based on descriptions by utilizing diffusion tech- niques. Using a multi-stage diffusion proce...
Mohammad Shahnawaz Shaikh· Journal of Intelligent Compu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.