We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as...
Fang-Yuan Tu, Xiang-Yue Zhang, Yi-Yi Cai et al.· 0 citations
Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they have already identified correctly. We distinguish perceptual failure, where relevant evidence is not recognized, from process failure, where recognized evidence fails to...
Yu-Yang Dai, Bo-Fei Huang, Hong-Bo Zhang et al.· 0 citations
The results show that the proposed SkelGen4D weakly supervised skeleton modeling matches or surpasses fully supervised baselines while scaling to diverse object categories for high-quality text-driven mesh animation.
Hao Feng, Zhi Zuo, Jia-Hui Pan et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.