Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded bench...
Hao Wang, Tao Yu, Liu-Zhou Zhang et al.· 0 citations
J-Access is proposed, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway, positioning J-Access as a model-level diagnostic for assessing residual susceptibility in unlearn...
Zirui Song, Hua-Xing Liu, Xiang Wang et al.· 1 citation
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must re...
Xin-Ye Li, Lingshuai Lin, Lei Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.