Skip to content
Preprint

Targeting the Attention Heads Behind Object Hallucination in LLaVA

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

A diagnosis-to-intervention pipeline for object hallucination is presented, and a controlled account of what acting on the diagnostic signal actually does is presented: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.

Abstract

Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure can guide a targeted fix, and we measure what that fix actually changes. We rank attention heads by how much their image attention drops around hallucinated object words, then screen the shortlist by ablating candidate heads and measuring the change in hallucination-token log probability, yielding a 32-head set. We restrict two interventions to these heads: a head-sliced LoRA adapter and an inference-time grounding controller. On 400 held-out COCO images, the combined method lowers CHAIRs (the fraction of captions with a hallucinated object) from 0.370 to 0.230 and CHAIRi (the fraction of hallucinated object mentions) from 0.156 to 0.096 (p<0.001, paired sign-flip tests). Two controls sharpen attribution. A random-head LoRA control, matched layer-for-layer and trained identically, performs no better than the matched baseline on a separate 200-image control split, supporting the role of head selection rather than LoRA capacity. Under fixed decoding budgets, the CHAIR reduction persists and grows with budget (23% at 64 tokens to 58% at 128), arguing against a pure max-token or truncation artifact, although the method remains shorter and more conservative. The resulting behavior reduces unsupported object mentions while also lowering object recall (0.78 to 0.70). We present a diagnosis-to-intervention pipeline for object hallucination, and, more importantly, a controlled account of what acting on the diagnostic signal actually does: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

ReWEIGH is a training-free decoding intervention that aggregates vocabulary ranks across visual positions and compares each candidate with a token-specific reference estimated from unlabeled images and applies a bounded penalty only to candidates that fall below their reference.

Jihae Jeong, Jun-Ha Choi, Hwanjo Yu · 0 citations
Open access Aug 2026

Mitigating Hallucination in Long Referring Expressions via Training-Free, Anchor-Preserved Visual Grounding

Long referring expressions create two coupled sources of hallucination in visual grounding. A detector can select an object that matches only part of the instruction, while a structured vision–language model (VLM) branch can hallucinate a target head or an attribute–object binding. We propose DeRecG, a training-free, a...

Hao-Xuan Song, Li-Huan Shao · 0 citations
#machine learning Preprint Aug 2026

VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs

VisER is proposed, a training-free two-sided metric for object-level hallucination detection that improves AUROC and AUPR over a range of baselines and measures whether object-context compatibility is backed by object-specific evidence from image tokens.

Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie et al. · 0 citations
#artificial intelligence Preprint Aug 2026

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

Multimodal large language models (MLLMs) fre- quently generate text that is inconsistent with the input image. While object- and attribute-level hallucinations have received considerable attention, relational hallucinations (incorrect de- scriptions of spatial or interactive relationships between objects) remain largel...

Siddhi Patil, N. Saxena, William B. Andreopoulos · 0 citations
Preprint Aug 2026

Test-Time Hallucination Control in Large Vision-Language Models

Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-world applications. Existing mitigation strategies can be categorized into training-based and training-free methods. Training-based methods...

Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian et al. · 0 citations
Preprint Aug 2026

Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM

Adversarial Contrastive Fine-Tuning (ACFT) uses an Adversarial Hallucination Attribute Flipping procedure, involving minimal, targeted adversarial perturbations that flip an image's hallucination attribute, to construct perfectly aligned positive-negative pairs, which are then used for contrastive fine- tuning.

Pei-Yang Xu, Xiaopei Zhu, Jun Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.