Jul 2026
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
CRISP is proposed, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context that serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.
Xu Li, Yi Zheng, Mengyang Zhao et al.
· arXiv.org · 1 citation