Skip to content

Author

Minjing Dong

We have 4 of 57 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as...

Yi-Cheng Xue, Han Wu, Ju-Feng Yang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly at...

Hao-Yang Luo, Lin-Wei Tao, Jie Gui et al. · 0 citations
Preprint Sep 2026

Region-Level Policy Optimization for Fine-grained MLLM Perception

Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...

Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al. · 0 citations
Preprint Aug 2026

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

RoRA is a training-free framework that casts visual token pruning as role-oriented regional evidence allocation, and consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios.

Qiyanhui Lu, Han Wu, Rong-Jia Xu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.