Zero-shot 3D visual grounding aims to localize specific objects based on textual descriptions and 3D visual input. However, the effectiveness of existing methods is significantly hindered by the ambiguous query text and deficient viewpoints. To address these issues, we propose TDVR, a training-free reasoning framework...
Qingxi Du, Junbo Wang, Yu-Ke Li et al.· 0 citations
Experimental results show that ToE substantially improves both problem-solving performance and efficiency and organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient...
Zi-Hao Deng, Yi-Ning Zhu, Lei-Ming Wang et al.· 0 citations
This work proposes a novel zero-shot video captioning framework (WSV) consisting of two training stages, which first generates corresponding synthetic video latent representations via a pretrained text-to-video generation model, and designs a polisher capable of bridging the gap between real and synthetic video distrib...
Liangyu Fu, Jun-Bo Wang, Yu-Ke Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.