Skip to content
Preprint

Multi-Faceted Evaluation and Mitigation of Emotion Hallucinations in MLLMs

Sep 2026 · 0 citations · 50 references
Computer Science

TL;DR

EHR (Emotion Hallucination Rate), an evaluator that quantifies emotion hallucinations across six facets, and HMER (Hallucination-aware Memory-guided Emotion Reasoning), a training-free framework for emotion hallucination mitigation, which enables fine-grained mitigation across diverse hallucination facets.

Abstract

Multimodal large language models (MLLMs) have shown strong potential in open-ended emotion understanding, yet they often generate emotion hallucinations. Evaluating such hallucinations is particularly challenging for two reasons. First, emotion understanding spans multiple cognitive facets, from multimodal perception to psychological reasoning. Second, emotional interpretations are expressed in free-form language, making existing closed-ended protocols insufficient for evaluation. To address these challenges, we introduce EHR (Emotion Hallucination Rate), an evaluator that quantifies emotion hallucinations across six facets: expression, action, audio, instinct, logic, and conclusion. Using EHR, we reveal that existing mitigation methods often reduce hallucinations in some facets while aggravating them in others, exposing the limitation of coarse-grained correction and the need for facet-aware localization and mitigation. Motivated by this finding, we propose HMER (Hallucination-aware Memory-guided Emotion Reasoning), a training-free framework for emotion hallucination mitigation. HMER maintains a Hallucination Memory that records localized hallucinated claims and enables targeted logit rectification, together with an Anchor Memory that preserves reliable intermediate reasoning states to stabilize subsequent generation. By selectively suppressing unreliable cues while preserving trustworthy reasoning context, HMER enables fine-grained mitigation across diverse hallucination facets. Extensive experiments on 19 MLLMs demonstrate the prevalence of emotion hallucinations and the effectiveness of our framework across diverse model architectures.

View source

Similar papers

Open access Sep 2026

Large Language Models Create Hallucinations in Response to Negated Text

Large language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negate...

Jaehyung Seo, Hyeonseok Moon, Heu-Jeoung Lim · 0 citations
Open access Sep 2026

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification...

Rashmi Nagpal, Unyimeabasi Usua, Kailey Simons et al. · 0 citations
Review Open access 2026

Extracting creativity from hallucination: Rethinking large language models

Hallucinations in large language models (LLMs) are always seen as limitations. However, could they also be a source of creativity? This survey explores this possibility, suggesting that hallucinations may contribute to LLM application by fostering creativity. Hallucinations are not treated as creativity perse; rather,...

Xuhui Jiang, Yi Liu, Ying-Han Shen et al. · 0 citations
#machine learning Preprint Sep 2026

MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models

This work introduces MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering two hallucination categories, and proposes a groundedness judge that uses reference rubrics and judge prompts guided by human annotations.

Wen-Soi Zhi, Giulio Segalini, Jian-Jia Chen et al. · 0 citations
Preprint Sep 2026

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

It is argued that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone, because improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities.

Mehrdad Fazli, Sina Mansouri, Mohit Marvania et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.