This work aggregates patch features from all views onto a single equirectangular panoramic canvas, and introduces a spatial pretraining curriculum by procedurally placing patch features of objects at chosen 3D world positions on an otherwise empty canvas, generating on-the-fly supervision spanning a broad range of spat...
Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner· arXiv.org· 0 citations
SpatialSpeak is introduced, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning and achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Yang Cao, Jia-Xin Zhang, Dave Zhenyu Chen et al.· 0 citations
This model matches optimization-based methods while delivering nearly 800x speedup, and surpasses the zero-shot performance of state-of-the-art generalizable models with markedly fewer parameters and reduced training/inference overhead, achieving an overall efficiency improvement.
Yingji Zhong, Dave Zhenyu Chen, Fu-Zhao Ou et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.