WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through ligh...