Jul 2026· International Conference on Machine Vision and Applications· Vol 14270, pp. 142700C - 142700C-10· 0 citations· 23 references
Engineering
TL;DR
This paper proposes π³-LEGS, an efficient geometric inference system designed for scalable long-sequence 3D reconstruction that maintains stable performance on thousand-frame sequences without runtime failures, highlighting its effectiveness in achieving a favorable accuracy efficiency trade-off for large-scale 3D reconstruction.
Abstract
In recent years, Transformer-based 3D vision foundation models have demonstrated strong generalization in multiview geometry and scene reconstruction. However, their scalability to long-sequence, urban-scale RGB streams remains limited due to the quadratic complexity of attention mechanisms, redundant frame processing, and high GPU memory pressure caused by dense spatial tokens. Although VGGT-Long partially alleviates these issues through chunked inference and loop-closure optimization, its geometric reasoning pipeline still involves substantial redundant computation, making it difficult to balance efficiency and accuracy in long-sequence scenarios. In this paper, we revisit the computational bottlenecks in long-sequence geometric inference, with a particular focus on spatial token redundancy and cross-frame global attention. We propose π³-LEGS, an efficient geometric inference system designed for scalable long-sequence 3D reconstruction. π³-LEGS integrates three key components: (1) a π³-based permutation-equivariant inference module to enhance unordered multi-view feature aggregation; (2) a geometry-aware keyframe selection mechanism that dynamically filters low-contribution frames to reduce redundant computation; and (3) a training-free block-sparse attention strategy that adaptively generates sparse attention masks based on pooled Query–Key similarity, significantly reducing global attention overhead. Extensive experiments on the KITTI Odometry dataset demonstrate that π³-LEGS achieves an average Absolute Trajectory Error (ATE) of 26.51, improving upon VGGT-Long by 6.4%, while reducing end-to-end inference time by 15.9%. Moreover, the proposed system maintains stable performance on thousand-frame sequences without runtime failures, highlighting its effectiveness in achieving a favorable accuracy efficiency trade-off for large-scale 3D reconstruction.
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information pr...
Jing-Ke Zhou, Chen-Hang Ma, Zhi-Zhou Zhong et al.· 1 citation
GeoWeaver, a unified framework comprising a Geometric Prior Model (GPM) and Test-Time Adaptation (TTA) and a robust CDF-style objective jointly optimizes weighted 2D reprojection and 3D consistency residuals, is presented.
Tinghao Jiang, Sheng Tang, Shengzhe Wei et al.· 0 citations
A depth-guided lightweight multi-task 3D scene understanding method based on a lightweight U-Net architecture, uses RGB and depth four-channel dual- modal input, and constructs a network structure with a shared encoder and independent decoders, allowing a single inference to simultaneously accomplish the three core tas...
Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks...
Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limite...
En-Peng Li, Yun-Zhou Zhang, Zhiyao Zhang et al.· 0 citations
Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplic...
Wei Zhang, Yihang Wu, Song Li et al.· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.