Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to...
Chenyangguang Zhang, Malgorzata Gwiazda, Guan-Long Jiao et al.· 0 citations
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close th...
Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.· 0 citations
Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.
Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al.· 1 citation
GenRec is introduced, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow, and attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved o...
Ata Çelen, Jaewoo Jung, Federico Tombari et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.