Skip to content

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

Jul 2026 · arXiv.org · Vol abs/2607.16326 · 1 citation · 29 references
Computer Science

TL;DR

CRISP is proposed, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context that serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.

Abstract

Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the cost of efficiency and noisy cross-modal attention. To address these limitations, we propose CRISP, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context. CRISP works in a two-stage pipeline: Stage 1 first identifies text-aligned visual tokens, and Stage 2 enhances contextual completeness through semantic diversity. Extensive experiments on LLaVA-1.5 and LLaVA-NeXT demonstrate that CRISP achieves superior performance retention under aggressive pruning ratios, maintaining up to 99.5% accuracy while reducing inference cost and latency by more than 2 times. CRISP serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens,...

Guang-Chuan Lv, Dian-Xing Shi, Ding-Jie Fu · 0 citations
#artificial intelligence Preprint Sep 2026

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measu...

Yi-Chen Guo, Tinghao Wang, Qizhe Zhang et al. · 0 citations
Preprint Aug 2026

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence...

Zizhong Ding, Junxian Li, Kai Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

SinkPruner is proposed, a training-free visual token pruning framework for efficient MLLM inference that follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains toke...

Shi-Yu Li, Zi-Yuan Hu, Shijia Huang et al. · 1 citation
#small language model Preprint Aug 2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP is proposed, a training-free progressive visual token pruning framework that removes redundant visual tokens based on the embedding similarity of input tokens before reasoning of the LLM backbone, and then prunes tokens during reasoning via head-aware pruning.

Chao-Fang Ma, Lin Jiang, Carol Jingyi Li et al. · 0 citations
Preprint Sep 2026

StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models

Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selecti...

Han-Sen Zhang, Lan He, Min Yao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.