Aug 2026· IEEE Transactions on Image Processing· Vol 35, pp. 8719-8733· 19 citations· ⚡ 2 influential· 48 references
MedicineComputer Science
Abstract
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. To further boost the performance, we advance the motion diffusion to motion-aligned temporal latent diffusion by developing a novel motion VAE. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style references. Text prompts are effective at defining semantic content, but they are often limited in capturing fine-grained style details such as timing, limb articulation, and...
Kai-Wei-Xian Lan, Bodie Criswell, Briana Fedkiw et al.· 0 citations
VISTA achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content-style balance at inference time.
Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek· 0 citations
Multi-modality motion stylization presents a solution to the challenge of generating flexible, stylized motion based on multimodal style inputs. Historically, motion stylization has grappled with the difficulty of balancing content and style, often prioritizing one at the expense of the other. This paper addresses the...
Wenfeng Song, Xingliang Jin, Shuai Li et al.· IEEE Transactions on Pattern...· 0 citations
Discrete text-to-motion generation models are powerful tools for synthesizing diverse human motions, yet endowing them with controllable, high-fidelity stylistic variations remains challenging. Direct retraining or fine-tuning these models often disrupts the fragile discrete token distributions learned by masked genera...
Shuai-Ying Hou, Hong-Yu Tao, Jun-Jie Gao et al.· IEEE Transactions on Visuali...· 0 citations
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requi...
Chen-Chieh Liao, Yi-Chen Peng, Yi-Yi Cai et al.· 0 citations
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bi...
Wan-Jiang Weng, Yong-Liang Wu, Xiaofeng Tan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.