Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection...
Fang Li, Jiraphon Yenphraphai, Quentin Herau et al.· 0 citations
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pa...
Wen-Bin Teng, Tian-Shuo Xu, De-Pu Meng et al.· 0 citations
ReWorld separates the two during training and bounds them at inference, and under a three-axis protocol covering action following, long-horizon recall, and video quality, against six recent interactive world models it attains the best control fidelity.