Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object de...
Cun-Zheng Fan, Dawei Yan, Guan-Lin Wang et al.· 0 citations
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios...
Zhen-Hao Shang, Hai-Zhao Jing, Hao-Kui Zhang et al.· 0 citations
This paper revisit VLM inference and presents a new efficient guidance scheme that complements similarity-based guidance, and proposes Cross Modal Residual (CMR), a training-free visual token compression method that combines CMR, text-attention relevance, and residual-space diversity to retain task-relevant and complem...
Congyang Ou, Rui-Ke Song, Yang Zhou et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.