Skip to content

Author

Shizhe Chen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Pretraining Body Part Representations for Text-Motion Retrieval

Text-motion retrieval has gained increasing research attention, yet several critical challenges remain such as data scarcity, limited fine-grained matching capabilities, and inadequate evaluation protocols. To address these issues, we propose POP-TMR which pretrains body part representations for fine-grained text-motion retrieval. Our approach leverages large-scale human motion datasets to pretrain a spatio-temporal transformer-based motion encoder, enabling more generalizable motion features. In addition to matching global motion and text representations, we propose a local branch to capture detailed body part features for enhancing spatial-aware cross-modal alignment. To improve evaluation, we introduce HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment. Extensive experiments on KIT-ML, HumanML3D and HumanML3D+ benchmarks demonstrate that POP-TMR outperforms state-of-the-art methods. Furthermore, we showcase its effectiveness in additional downstream applications, including text-to-motion generation evaluation, human interaction recognition and zero-shot moment retrieval. Data, code and pretrained model are publicly available at https://lin-kayla.github.io/POP_TMR/.

Kejun Lin, Shizhe Chen, Anwen Hu et al. · 0 citations