A novel framework for motion blending across heterogeneous skeletons is presented, which combines a semantic encoder, which extracts per-frame latent representations of the motion state, with a diffusion-based decoder, which reconstructs character-specific motion conditioned on this latent code.
Abstract
Motion blending in character animation enables the synthesis of new motions by interpolating between existing examples. Current methods are typically restricted to fixed skeleton topologies, requiring identical or near-identical skeletal structures across characters. We present a novel framework for motion blending across heterogeneous skeletons. The proposed architecture combines a semantic encoder, which extracts per-frame latent representations of the motion state, with a diffusion-based decoder, which reconstructs character-specific motion conditioned on this latent code. At inference, blended motions are obtained by interpolating the latent representations of two input motions. We train and evaluate the method on the Truebones Zoo dataset using motions defined on both same and distinct skeleton topologies, demonstrating the ability to achieve smooth and plausible blending in a variety of scenarios.
This work introduces a framework for cross-morphology motion transfer with semantic style alignment that uses morphology-agnostic control signals (e.g., velocity, angular velocity, relative height) to align behaviors across species.
Alexios Mylordos, J. L. Pontón, Nuria Pelechano et al.· IEEE Transactions on Visuali...· 0 citations
UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.
Lin-Zhan Mou, Jiahui Lei, Zhi-Yang Dou et al.· 0 citations
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent anim...
Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu et al.· 0 citations
AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements and enhances usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance.
Kwan Yun, Serin Yoon, Sun-Jin Jung et al.· IEEE Transactions on Visuali...· 0 citations
Wan-Animate-2 is presented, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer and achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely.
Guangyuan Wang, Liucheng Hu, Dechao Meng et al.· 2 citations
ViP-Rig is a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones into a frozen pretrained autoregressive generator.
Zihan Qin, Ming-Ze Sun, Yifan Mao et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.