Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection...
Fang Li, Jiraphon Yenphraphai, Quentin Herau et al.· 0 citations
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pa...
Wen-Bin Teng, Tian-Shuo Xu, De-Pu Meng et al.· 0 citations
TerraZero is a procedural driving simulator and self-play training stack that meets the goals of reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.
Zhouchonghao Wu, Akshay Rangesh, Weixin Li et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.