Skip to content

Pretraining Body Part Representations for Text-Motion Retrieval

Jul 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 65 references

Abstract

Text-motion retrieval has gained increasing research attention, yet several critical challenges remain such as data scarcity, limited fine-grained matching capabilities, and inadequate evaluation protocols. To address these issues, we propose POP-TMR which pretrains body part representations for fine-grained text-motion retrieval. Our approach leverages large-scale human motion datasets to pretrain a spatio-temporal transformer-based motion encoder, enabling more generalizable motion features. In addition to matching global motion and text representations, we propose a local branch to capture detailed body part features for enhancing spatial-aware cross-modal alignment. To improve evaluation, we introduce HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment. Extensive experiments on KIT-ML, HumanML3D and HumanML3D+ benchmarks demonstrate that POP-TMR outperforms state-of-the-art methods. Furthermore, we showcase its effectiveness in additional downstream applications, including text-to-motion generation evaluation, human interaction recognition and zero-shot moment retrieval. Data, code and pretrained model are publicly available at https://lin-kayla.github.io/POP_TMR/.

View source