Skip to content

Author

Haokui Zhang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object de...

Cun-Zheng Fan, Dawei Yan, Guan-Lin Wang et al. · 0 citations
Preprint Sep 2026

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios...

Zhen-Hao Shang, Hai-Zhao Jing, Hao-Kui Zhang et al. · 0 citations
Preprint Aug 2026

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

This paper revisit VLM inference and presents a new efficient guidance scheme that complements similarity-based guidance, and proposes Cross Modal Residual (CMR), a training-free visual token compression method that combines CMR, text-attention relevance, and residual-space diversity to retain task-relevant and complem...

Congyang Ou, Rui-Ke Song, Yang Zhou et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.