Skip to content

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

Sep 2026 · 0 citations
Computer Science

TL;DR

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Abstract

Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens. Beyond token reduction, recent studies exploit vertical redundancy through early exit or fixed-layer skipping. However, we find that the extent and distribution of this redundancy vary across inputs and differ between self-attention and MLP modules. Motivated by these observations, we propose AdaVSkip, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules. These decisions collectively define an input-specific visual-computation path, but their discrete and non-differentiable nature makes learning effective paths challenging. To address this challenge, we develop a progressive two-stage training framework that updates only the routers while keeping the backbone frozen. Stage I establishes an initial routing policy through supervised training with input-specific targets derived from module-wise necessity scores. To further align the routing policy with task performance, Stage II uses reinforcement learning to optimize routing decisions with direct feedback from generated answers. It combines an answer correctness reward with a skip-consistency reward that discourages excessive retention of visual-token computation. Across three MLLM backbones, AdaVSkip maintains strong task performance with substantially less computation. On LLaVA-NeXT-7B, AdaVSkip reduces FLOPs by 53.2\% while preserving the original model's average performance. Combining it with visual token compression increases this reduction to 91.2\%, while retaining 97.2\% of the original performance on average.

View source

Similar papers

Preprint Sep 2026

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

STD is proposed, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases.

Shuo Zhang, Jin-Tao Tong, Yi-Xiong Zou et al. · 0 citations
Preprint Sep 2026

StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models

Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selecti...

Han-Sen Zhang, Lan He, Min Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whe...

Jing-Di Lei, Junxian Li, Di Zhang et al. · 0 citations
#computer vision Preprint Aug 2026

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning

Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance f...

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

SinkPruner is proposed, a training-free visual token pruning framework for efficient MLLM inference that follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains toke...

Shi-Yu Li, Zi-Yuan Hu, Shijia Huang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco...

Hao-Yu Guo, Yuan Feng, Junlin Lv et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.