Skip to content

TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

Jul 2026 · arXiv.org · Vol abs/2607.04593 · 0 citations · 36 references
Computer Science

TL;DR

This work introduces TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model.

Abstract

Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction methods rely on attention-based scores or pairwise similarity, without an explicit semantic representation of each token. We introduce TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model. TORINO leverages Sparse Autoencoders (SAEs) to project visual tokens into an interpretable latent space where token relationships can be analyzed through shared concept activations. Specifically, we define concept overlap as the degree of agreement between active SAE latents and use it to group tokens that share semantic content. Reduction within each group is then performed by either pruning or merging, providing a unified framework that preserves semantically important visual information while removing redundancy. Unlike fixed-budget approaches, TORINO dynamically adapts the reduction rate to input complexity, allowing different images to retain different numbers of tokens. Experiments across multiple vision-language benchmarks show that TORINO achieves favorable efficiency-accuracy trade-offs, reducing the number of visual tokens with minimal performance loss.

View source

Similar papers

Jul 2026

GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

This work proposes Greedy Orthogonal Token Selection (GOTS), a training-free and query-agnostic method that achieves higher average performance retention than the strongest evaluated baselines, and a controlled OCRBench study shows that it reduces model-side time-to-first-token after accounting for selection overhead.

Jun Ling, Tao Huang, Junzhuo Liu et al. · 0 citations
Preprint Aug 2026

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

The results show that structured coordinate generation provides an effective approach to generative visual grounding and Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM arch...

Xiuyuan Zhu, Ke Lu, Kun Dong et al. · 0 citations
Jul 2026

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

Despite-encoder vision-language models expose a similarity interface that enables zero-shot retrieval but fails compositional constraints, this work proposes factored inference, which separates evidence extraction from constraint execution, and introduces LCSE (Logic-Constrained Score Editing), a training-free method t...

S. Alshehri, Zhan-Tao Yang, Han Zhang et al. · 0 citations
#computer vision Preprint Sep 2026

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Image tokenizers define the ``visual language''of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed...

Siting Li, Zheng-Yang Wang, S. Du et al. · 0 citations
Preprint Aug 2026

Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

Reducing the representation within retained tokens provides an effective complement to token pruning for aggressive VLA compression, and can be applied to language values, allowing visual and language representations to be compressed without removing additional tokens.

Wei Jiang, Wei Wang · 0 citations
Aug 2026

Instructing the Learning of Language Model with the Token Interpretation to Improve Language Understanding

Pretrained language models (PLMs) have established state-of-the-art performance across diverse natural language understanding (NLU) tasks. This study reveals that seman-tic-rich explanations of lexical units can effectively guide PLM learning processes. We propose a novel language understanding enhancement method with...

Tianyi Chen, Yashen Wang, Huan Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.