It is argued that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone, because improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities.
Abstract
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.
Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing ben...
When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by asking: can we use large vision-language models (LVLMs) as experimental instruments for studying their own failure dynam...
Zhi-Peng Zhao, Wen-Xu Wang, Peishun Liu et al.· 0 citations
This work introduces MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering two hallucination categories, and proposes a groundedness judge that uses reference rubrics and judge prompts guided by human annotations.
Wen-Soi Zhi, Giulio Segalini, Jian-Jia Chen et al.· 0 citations
ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.
Ze-Xing Zhang, Tian-Yang Lei, Ke-Wei Yang et al.· Proceedings of the 32nd ACM...· 0 citations
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence, so a fully black-box framework that models hallucination as a structured uncertainty pattern is proposed.
Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia et al.· 0 citations
EHR (Emotion Hallucination Rate), an evaluator that quantifies emotion hallucinations across six facets, and HMER (Hallucination-aware Memory-guided Emotion Reasoning), a training-free framework for emotion hallucination mitigation, which enables fine-grained mitigation across diverse hallucination facets.
Bo-Wen Zeng, Pei-Pei Song, Wei-Dong Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.