Accurate Surgical Scene Reconstruction from Multi-view FoundationStereo
Abstract
Accurate 3D reconstruction of surgical scenes is a critical enabling technology for advancements in intraoperative navigation, surgical training, and robotic automation. While learning-based stereo depth estimation methods have demonstrated high in-domain accuracy, their performance often degrades significantly under domain shifts. This limitation is particularly acute in surgical applications, where large-scale, annotated datasets are scarce. In this work, we investigate the application of FoundationStereo, a recently proposed vision foundation model for stereo matching, to the task of surgical scene reconstruction. We leverage its zero-shot, single-frame depth estimation capabilities within a multi-view fusion framework based on the Truncated Signed Distance Function (TSDF) to achieve comprehensive scene reconstruction. Our experiments, conducted on the public SCARED dataset captured with a da Vinci Xi surgical robot, demonstrate that FoundationStereo achieves state-of-the-art zero-shot accuracy. We report a sub-millimeter mean error for single-frame depth estimation and an error under 2 mm for fused multi-view reconstructions, significantly outperforming the zero-shot STTR baseline. These results highlight the substantial potential of FoundationStereo for enabling accurate, high-fidelity surgical scene reconstruction without domain-specific training. We also discuss current limitations, including the reliance on known camera pose information and challenges in dynamic scenes, and outline future research directions to enhance robustness for clinical applications.