Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed...
J. Shi, Kai-Chen Zhou, Haoyu Chen et al.· 0 citations
This work presents a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals.
Run-Jia Zeng, Hang Hua, Yiyang Liu et al.· 0 citations
VideoArgus is introduced, a unified rubric-grounded framework covering five video generation and editing settings that achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks.
Ziyun Zeng, Zi-Xuan Wang, Yongsheng Yu et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.