Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bi...
Wan-Jiang Weng, Yong-Liang Wu, Xiaofeng Tan et al.· 0 citations
Each task and its evaluation protocol is described, the challenge leaderboards are presented, and the leading submissions are summarized, with the aim of documenting the current state of each task as measured on held-out challenge data.
A. Cioppa, Silvio Giancola, Haakan Ardo et al.· 1 citation
A comprehensive external and internal investigation of multimodal in-context learning on the image captioning task is conducted, revealing both how ICEs configuration strategies impact model performance through external experiments and characteristic typical patterns through internal inspection.
Li Li, Yongliang Wu, Jingze Zhu et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.