Aug 2026· International Journal of Computer Vision· Vol 134· 0 citations· 83 references
TL;DR
This work proposes REVERIE+ (ReflEctiVERatIonalE), an extension of REVERIE substantially expanded in domain diversity, task complexity, and annotation richness, tailored to advanced LVLMs, which broadens domain coverage and increases task difficulty, while improving annotation reliability.
This work proposes a training-free hallucination mitigation framework for dynamic, per-instance suppression at test time, and proposes a dynamically combined projection that selectively suppresses the most probable hallucination directions while preserving image-grounded semantics.
Ali Cheraghian, Hamidreza Dastmalchi, Hamed Barzamini et al.· 0 citations
CAL-RAG (Context-Aware Low-Rank Calibration for RAG), a parameter-efficient fine-tuning and decoding calibration framework designed to enforce strict contextual faithfulness without compromising generative fluency, is proposed.
Sophia N. Tawar, Liam K. Peing, Amani Bellow· International Journal of App...· 0 citations
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with...
Zian Ding, Zi-Lin Zhao, Ying-Jie He et al.· 0 citations
LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens, consistently improves Best-of-$N$ selection over existing baselines with negligible additional overhead.
Beomsik Cho, Jin-Ha Kim, Dongseok Lee et al.· 0 citations
TruthShield is presented, a metric-aware trigger-guided QLoRA adapter training and evaluation pipeline for hallucination-aware language model adaptation and suggests that trigger-guided adapter training may learn surface-level response patterns without clear evidence of semantic hallucination mitigation under the curre...
EviAnchor is proposed, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation in large vision-language models and demonstrates consistent improvements in visual grounding.
Sihang Jia, Shuliang Liu, Song-Bo Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.