Skip to content

Two2Four: Generative Quadruped Puppeteering from Human Motion

Jul 2026 · Computer graphics forum (Print) · Vol 45 · 0 citations · 43 references
Computer Science

TL;DR

An automatic human-to-quadruped puppeteering framework that produces plausible and controllable quadruped motions from ordinary human motion data and enables fine-grained intuitive control of the quadruped motion such as head movement control and individual limb puppeteering.

Abstract

Realistic animal motion for virtual production is typically obtained either through motion capture of highly trained performers who accurately mimic animal behavior, or by retargeting ordinary human motion using complex control setups. Both approaches are challenging and often fail to fully reproduce the nuances of natural animal motion, motivating data-driven alternatives. We present an automatic human-to-quadruped puppeteering framework that produces plausible and controllable quadruped motions from ordinary human motion data. Our approach employs a two-stage generative diffusion model trained purely on quadruped motion data. By introducing a structured conditioning and inpainting strategy, our method supports a wide range of actions, including walking, running, jumping, sitting, and lying. Furthermore, we enable fine-grained intuitive control of the quadruped motion such as head movement control and individual limb puppeteering. Experimental results demonstrate improved motion realism and controllability compared to existing retargeting approaches, highlighting the effectiveness of our framework as a tool for animation and virtual production applications.

View source

Similar papers

Preprint Sep 2026

Kirin: Animal Motion Generation from In-the-Wild Video

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion can be captured in controlled environments, it is impractical for most animal species, resulting in...

Brian Nlong Zhao, Zhuoyang Pan, J. Rehg et al. · 0 citations
Preprint Sep 2026

UniMo: Unifying Human and Animal Motion Generation

The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending recent advances in text-driven human motion generation to the animal domain remains challenging due to two core limitations. First, animals e...

Ze-Yu Zhang, Zhi-Yuan Zhang, Siheng Wang et al. · 0 citations
Preprint Sep 2026

MotionCanvas: Learning Implicit Motion Planning from Composable Kinematic Cues

Professional character animation requires both natural motion and precise, versatile control. For example, it is common for the creators to define the timing of a specified action, to control the motion range of the character's arm swing, and the route the character walks through, like specifying various kinematic moti...

Ze-Yu Ling, Di Kang, Qing Shuai et al. · 0 citations
Preprint Aug 2026

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

This work introduces HumanTracker, a preference-aligned metric trained on 12K motion pairs containing 24K motions that better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.

Dai-En Liu, Ze-Kun Qi, Jia-Yu Zeng et al. · 0 citations
Preprint Aug 2026

AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures

For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles...

Guoxing Sun, Heming Zhu, Linjie Lyu et al. · 0 citations
Preprint Sep 2026

Multi-Modal Controlled Coherent Motion Generation

This paper proposes MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data.

Yi-Fei Liu, Qiong Cao, Hong-Wei Yi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.