ShapeRhythm: Personalized and Harmonic Dance Generation Bridging Diverse Bodies and Melodies
Abstract
Deriving vivid and expressive 3D human dance from music signals has gained tremendous progress in virtual avatar animation. Existing approaches typically overlook the diversity of human body shapes and generate dance motions using standardized body representations, without explicitly modeling the intrinsic relationship between body shape variations (i.e., thinness or plumpness) and musical rhythm. As a result, the synthesized dance motion lacks both personalization and harmony. Moreover, the absence of high-quality human dance datasets, particularly those incorporating bodyshape variations and annotations, also poses a significant challenge for this problem. To address this issue, we first newly compile a large-scale 3D human dance dataset, named ShapeDancer, which encompasses 13,574 s of data paired with detailed body shape variations. Additionally, we propose ShapeRhythm, a novel framework that allows personalized and harmonic dance generation, bridging diverse human bodies with expressive melodies. Our ShapeRhythm framework is built upon unifying the music signals and shape-aware text descriptions as conditions. Specifically, to enhance the dance personalization, we propose FreeShape, a delicate shape similarity matching approach which realizes freestyle and open-vocabulary text-guided bodyshape conversion. Furthermore, we design FreeDance to diversify and harmonize the dance motion generation, while preserving the shape-aware temporal consistency with music rhythm. Meanwhile, a fine-grained physics-guided Foot-Ground Interaction Constraint (FGI) is introduced to alleviate foot-skating artifacts and enhance detailed foot rhythm coordination. Extensive experiments conducted on our newly collected ShapeDancer dataset demonstrate that ShapeRhythm achieves state-of-the-art performance, realizing personalized and harmonic dance generation. More details are available on our project page.