Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bi...
Wan-Jiang Weng, Yong-Liang Wu, Xiaofeng Tan et al.· 0 citations
Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity is proposed, which demonstrates that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base mo...
Jiawei Feng, Jiancan Wu, Xingyu Zhu et al.· 1 citation
This work proposes RITA, a Robust test-tIme prompt-TAdaptation framework that shifts from sample-level estimates to distribution-level alignment, and employs optimal transport to align the distribution of augmented visual features with textual prototypes, mitigating adversarial outliers and rectifying cross-modal seman...
Xingyu Zhu, Huanshen Wu, Shuo Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.