Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inf...
Han-Yuan Xiao, Gong-Lin Chen, Hao-Lin Xiong et al.· 0 citations
XDG is introduced, an efficient visual disambiguation model designed for scalable SfM that remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup.
Gong-Lin Chen, Ben Southall, Han-Yuan Xiao et al.· 0 citations
WA-JEPA is presented, a V-JEPA-native world-action model designed for autonomous driving planning that employs hybrid future-masked pre-training and recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downst...
Xin-Lin Wang, Yu Xiang, Yuheng Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.