Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant tem...
Kang He, Wen-Shuo Peng, Zi-Hui Gao et al.· 2 citations· ⚡1
GROVE is introduced, a training-free framework that supports both behaviors with one memory grown causally from a continuous video stream, and achieves the best results among the compared methods.
Sitong Gong, Caixin Kang, Tianyu Yan et al.· 0 citations
This work explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance, establishing two properties of Marionette, a world model for interactive games with articulated characters that is directly controllable.
Zian Meng, Zhen Li, Chuanhao Li et al.· 3 citations
Alaya-EVOKE (Evoke) addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation, and achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.