Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pa...
Wen-Bin Teng, Tian-Shuo Xu, De-Pu Meng et al.· 0 citations
Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inf...
Han-Yuan Xiao, Gong-Lin Chen, Hao-Lin Xiong et al.· 0 citations
XDG is introduced, an efficient visual disambiguation model designed for scalable SfM that remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup.
Gong-Lin Chen, Ben Southall, Han-Yuan Xiao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.