Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-conditioned video generation....
Rui-Qi Li, Xuan-Yi Liu, Si-Jia Li et al.· 0 citations
TRWORLDBENCH is introduced, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality.
Xuan-Yi Liu, Hao-Feng Wang, Rui-Qi Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.