Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking...
Xiaoqiang Lu, Li-Cheng Jiao, Ling-Ling Li et al.· 0 citations
Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce tas...
Xiaoqiang Lu, Licheng Jiao, Long Sun et al.· 0 citations
To relieve the problems of the lack of semantic alignment in the geometric representation space and insufficient modality-specific instruction calibration in infrared-visible image fusion (IVIF), we first propose a dynamic instruction-aware geometric representation (DIGR) for IVIF, which achieves semantic alignment bet...
Mengru Ma, Yong-Zhe Wang, Wen-Ping Ma et al.· IEEE Transactions on Geoscie...· 0 citations
Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method...
Recently, the zero-shot image captioning (zero-shot IC) method based on pre-trained visual language models (VLMs) and large language models (LLMs) has made significant progress. However, how to adapt it to the zero-shot video captioning (zero-shot VC) scenario (without video-text paired supervision) has not been well e...
Qianyue Bao, Fang Liu, Licheng Jiao et al.· IEEE Transactions on Image P...· 0 citations
A five-dimensional analytical framework is introduced that clarifies how EI transforms isolated search trajectories into cumulative scientific insight, and identifies critical bottlenecks regarding evaluation, process traceability, and shared infrastructure.
Chao Wang, Lingling Li, Fang Liu et al.· 0 citations