Jul 2026
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
This work introduces TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model.
Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda et al.
· arXiv.org · 0 citations