RVSD is proposed, a training-free and plug-and-play decoding framework that unifies token sparsification and Semantic-Space Visual Retrieval (SSVR) within a single decoding pass, and asemantics-directed token selection is introduced within RVSD, which selectively sparsifies redundant tokens while preserving critical visual information.
Abstract
Large vision-language models have achieved remarkable success in vision-language tasks. However, they remain prone to Visual Hallucinations (VHs), undermining their reliability in real-world applications. Existing solutions typically require curated datasets, additional training, or multi-round decoding, resulting in considerable computational overhead. In this paper, we propose \textbf{RVSD} (\underline{R}etrieval \underline{V}ision \underline{S}parse \underline{D}ecoding), a training-free and plug-and-play decoding framework that, for the first time, unifies token sparsification and \textbf{Semantic-Space Visual Retrieval} (SSVR) within a single decoding pass. Within RVSD, we introduce a \textbf{semantics-directed token selection} strategy that selectively sparsifies redundant tokens while preserving critical visual information. We further propose the SSVR mechanism, which reformulates visual compensation as an on-demand cross-modal retrieval process within a shared semantic space. Extensive experiments demonstrate that RVSD achieves state-of-the-art performance in mitigating VHs while maintaining robust suppression capabilities under long-context generation settings. Our code is available here.\footnote{https://github.com/canjie-liu/RVSD}
UniProbe is introduced, a lightweight, unified, learnable detector that models a frozen LVLM's heterogeneous computational trace from a single forward pass and achieves state-of-the-art token-level and object-hallucination detection.
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca et al.· 0 citations
EviAnchor is proposed, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation in large vision-language models and demonstrates consistent improvements in visual grounding.
Sihang Jia, Shuliang Liu, Song-Bo Yang et al.· 0 citations
The framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives.
Aditi Sarker, Rafi Ibn Sultan, Hui Zhu et al.· 0 citations
This time, the SHROOM-Visions task aims to tackle hallucinations through a model-agnostic detection task focused on large vision-language models, building on the recently introduced SHEEP dataset.
Raúl Vázquez, Aman Sinha, Chuyuan Li et al.· 0 citations
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-world applications. Existing mitigation strategies can be categorized into training-based and training-free methods. Training-based methods...
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian et al.· 0 citations
SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.