Skip to content
#computer vision Preprint

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features, is proposed.

Abstract

Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the appropriate middle layer requires processing numerous visual tokens through several language model layers, by which point considerable computation has already been spent. To address both problems, we propose Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features. During inference, MAP combines the predicted importance scores with a diversity criterion to prune visual tokens before the first language model layer. Thus, MAP requires no attention maps for pruning and remains compatible with existing inference acceleration techniques. Across ten benchmarks on LLaVA-NeXT-7B, MAP retains 97.5% of the unpruned model performance with only 5.56% of the visual tokens, yielding a 3.09x end-to-end speedup.

View source

Similar papers

Preprint Sep 2026

StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models

Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selecti...

Han-Sen Zhang, Lan He, Min Yao et al. · 0 citations
#small language model Preprint Aug 2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP is proposed, a training-free progressive visual token pruning framework that removes redundant visual tokens based on the embedding similarity of input tokens before reasoning of the LLM backbone, and then prunes tokens during reasoning via head-aware pruning.

Chao-Fang Ma, Lin Jiang, Carol Jingyi Li et al. · 0 citations
Jul 2026

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

Trend-aware Pruning is proposed, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem, and enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, the...

Jie Ma, Zhike Qiu, Jie Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
Preprint Aug 2026

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

A spatial novelty constraint is introduced that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas and prevents the retained tokens from concentrating in a few locally salient areas in E2S-Pruner.

Taoyu Qian, Qi Wang, D. Shi et al. · 1 citation
Preprint Sep 2026

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

STD is proposed, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases.

Shuo Zhang, Jin-Tao Tong, Yi-Xiong Zou et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.