Aug 2026· IEEE Transactions on Pattern Analysis and Machine Intelligence· Vol PP· 0 citations
Medicine
Abstract
Multi-modality motion stylization presents a solution to the challenge of generating flexible, stylized motion based on multimodal style inputs. Historically, motion stylization has grappled with the difficulty of balancing content and style, often prioritizing one at the expense of the other. This paper addresses the complex challenge of multi-modality content-style duality, achieving a sophisticated integration that both preserves and enhances the core narrative through nuanced stylistic modifications. We propose a Multi-modality Latent Diffusion Model (MM-LDM), a novel framework that leverages diffusion models under multi-modality conditions, including motion style (text-based or motion-based), motion content, and motion trajectory components. A central innovation in our approach is the introduction of a Multi-condition Denoiser, which carefully balances the preservation of primary content with the dynamic integration of style and trajectory as secondary conditions. This multi-modality guidance mechanism, implemented during the denoising process, ensures that new styles are seamlessly integrated with the original content. It gives rise to more authentic and cohesive motion stylization outcomes, establishing a new benchmark in computer animation. To further refine the control over the text-based motion style, we introduce an LLM parser that converts broad motion descriptions into detailed, part-specific representations. By decomposing the human body into movement-related parts, our method significantly enhances the precision and effectiveness of text-based motion stylization, enabling fine-grained control over individual body parts. Our model's effectiveness and generalization capabilities have been rigorously validated through extensive experiments, including text-based motion stylization and generating stylized motion with video sources, which all demonstrate the potential of our MM-LDM to advance the state-of-the-art motion stylization.
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differen...
Zhe Li, Yi-Sheng He, Lei Zhong et al.· IEEE Transactions on Image P...· 19 citations· ⚡2
VISTA achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content-style balance at inference time.
Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek· 0 citations
We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style references. Text prompts are effective at defining semantic content, but they are often limited in capturing fine-grained style details such as timing, limb articulation, and...
Kai-Wei-Xian Lan, Bodie Criswell, Briana Fedkiw et al.· 0 citations
Discrete text-to-motion generation models are powerful tools for synthesizing diverse human motions, yet endowing them with controllable, high-fidelity stylistic variations remains challenging. Direct retraining or fine-tuning these models often disrupts the fragile discrete token distributions learned by masked genera...
Shuai-Ying Hou, Hong-Yu Tao, Jun-Jie Gao et al.· IEEE Transactions on Visuali...· 0 citations
Introduction Multi-modal conditioning in latent diffusion models—combining text, structural, and spatial guidance signals—substantially improves controllable image synthesis, yet two limitations persist across most existing frameworks. First, conditioning modalities are fused using fixed architectural weights that rema...
S. Remya, Manu J. Pillai, Laveena Herman et al.· Frontiers in Artificial Inte...· 0 citations
This work introduces stylized phase manifolds—a compact, interpretable latent representation that disentangles motion content, the temporal structure, and style and develops a diffusion‐based motion generator that enables fine‐grained control over semantic, temporal, and stylistic aspects of motion.
Jing-Yuan Li, Peizhuo Li, A. Aristidou et al.· Computer graphics forum (Pri...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.