Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce tas...
Xiaoqiang Lu, Licheng Jiao, Long Sun et al.· 0 citations
To relieve the problems of the lack of semantic alignment in the geometric representation space and insufficient modality-specific instruction calibration in infrared-visible image fusion (IVIF), we first propose a dynamic instruction-aware geometric representation (DIGR) for IVIF, which achieves semantic alignment bet...
Mengru Ma, Yong-Zhe Wang, Wen-Ping Ma et al.· IEEE Transactions on Geoscie...· 0 citations
Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method...
Pu Luo, Cong Xu, Yu-Mei Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.