Skip to content

VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

VisLens (Visual Focus via Logit Lens), a Visual Search method built on the logit lens, which decodes the semantics held in a hidden state by projecting it through the LLM head, and which matches or exceeds prior baselines while delivering a substantial latency advantage.

Abstract

Multimodal large language models (MLLMs) struggle with fine-grained Visual Search, the task of locating small or rare objects in high-resolution images. Existing remedies fall into two families: (1) Training-free methods based on attention or confidence scores are accurate but slow, since they require multiple MLLM queries per example. (2) Reinforcement Learning (RL) trained tool-use models are faster at inference but opaque, since their tool calls remain uncontrollable and hard to interpret. To overcome this, we propose \emph{VisLens} (Visual Focus via Logit Lens), a Visual Search method built on the logit lens, which decodes the semantics held in a hidden state by projecting it through the LLM head. VisLens further uses a lightweight tuned-lens that maps early hidden states into the final hidden state space, so visual tokens can be read out from early layers. These tokens are matched to target words in the query to generate a crop of the relevant region, which is fed back in alongside the original image to produce the final answer. The whole process, from decoding to the final answer, completes in a single forward pass without repeated queries. VisLens matches or exceeds prior baselines while delivering a substantial latency advantage, running $8.5$--$9.9\times$ faster than Thyme and up to $22.2\times$ faster than training-free multi-pass search methods.

View source

Similar papers

Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al. · 2 citations
#artificial intelligence Preprint Sep 2026

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

SinkPruner is proposed, a training-free visual token pruning framework for efficient MLLM inference that follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains toke...

Shi-Yu Li, Zi-Yuan Hu, Shijia Huang et al. · 1 citation
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
#artificial intelligence Preprint Sep 2026

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the corre...

Zhen-Bin Wang, Lei Zhang, Li-Tuan Wang et al. · 0 citations
Preprint Sep 2026

StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models

Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selecti...

Han-Sen Zhang, Lan He, Min Yao et al. · 0 citations
Preprint Aug 2026

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

UniProbe is introduced, a lightweight, unified, learnable detector that models a frozen LVLM's heterogeneous computational trace from a single forward pass and achieves state-of-the-art token-level and object-hallucination detection.

Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.