Skip to content
Book Open access

Motion4Motion: Motion Transfer Across Subjects at Inference

Jul 2026 · International Conference on Computer Graphics and Interactive Techniques · 0 citations · 60 references
Computer Science

Abstract

This work explores the motion transfer from one video to another, which is crucial in animation for diverse characters. Previously, video motion transfer has been largely explored between human and human-like characters, enabling a lot of applications in digital creation. However, these approaches encounter a main limitation. Specifically, related technical pipelines heavily rely on a predefined human skeleton structure and accordingly require skeleton-conditional model training. On the one hand, these methods are difficult to generalize to diverse characters, such as animals from different species, while preserving their unique motion styles. On the other hand, labeled data in diverse skeletons is limited, which additionally restricts the large-scale training for the task. In this paper, we jump out of the skeleton-based motion transfer framework and propose a training-free motion transfer framework, named Motion4Motion. Motion4Motion models the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier. Extensive experimental results and novel applications show our methods outperform baselines impressively.

Read PDF

Similar papers

Preprint Aug 2026

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill-defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two-stage framework. Stage~I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference-video-conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT-Dataset and OpenVMT-Bench for training and evaluating image- and text-conditioned motion transfer across Same, Near, and Far category gaps. Extensive experiments demonstrate state-of-the-art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/

Zhixue Fang, Zhimin Zhang, Bi'an Du et al. · 0 citations
Open access Jul 2026

DMP: Directable Motion Retargeting through Motion Paraphrasing

Creating diverse and realistic human motions is a fundamental cornerstone of computer animation, with numerous applications in games, movies, and AR/VR. While motion capture is a valuable tool for capturing motions across varied body sizes, obtaining unique motion data for a variety of characters is often prohibitively expensive. Motion retargeting addresses this limitation by adapting existing motions to different character morphologies, however, existing approaches often involve trade-offs between motion realism, user control, and adaptability to artistic needs. In this work, we propose Directable Motion Paraphrasing (DMP), a novel motion retargeting framework based on the concept of motion paraphrasing, analogous to text paraphrasing, where the core semantics of a motion are preserved while allowing expressive, user-directed variations. Our framework constructs a large-scale motion paraphrasing dataset, which captures the diversity of human motion across different body shapes, and trains a diffusion-based generative model that learns both invariances and variations in motion. To enable user control during inference, we introduce a flexible mechanism for specifying spatio-temporal constraints, such as joint positions, rotations, and object interactions, which can be incorporated into the generative process through masked inpainting and loss guidance. We demonstrate the effectiveness of our framework through various examples, showing its ability to produce realistic, diverse, and controllable retargeted motions that meet the artistic demands of animation pipelines. Extensive experiments demonstrate the system’s flexibility, motion plausibility, and directability, highlighting its potential as a tool for intuitive and high-quality motion retargeting.

Sunmin Lee, Davis Rempe, Yifeng Jiang et al. · 0 citations
Preprint Jul 2026

MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer

It is demonstrated that MoSAIC improves the response--preservation trade-off required for selective and controllable part-local motion editing, and is presented as a latent diffusion framework for part-local reference-conditioned motion style transfer.

N. Amini, Kevin Desai · 0 citations
Preprint Jul 2026

MIME: Multimodal Interactive Motion Encoder

Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.

Addison Zucek, Prerit Gupta, Kamila Kuatova et al. · 0 citations
Jul 2026

Motion-driven 4D scene generation

Guo-Wei Yang, Qun-Ce Xu, Zhao Wei et al. · 0 citations