Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as...
Yi-Cheng Xue, Han Wu, Ju-Feng Yang et al.· 0 citations
Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly at...
Hao-Yang Luo, Lin-Wei Tao, Jie Gui et al.· 0 citations
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...
Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al.· 0 citations
RoRA is a training-free framework that casts visual token pruning as role-oriented regional evidence allocation, and consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios.
Qiyanhui Lu, Han Wu, Rong-Jia Xu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.