Skip to content

Reliability Challenges in Diffusion Vision-Language Models

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

dLVLMs reverse the yes-bias of AR models in binary visual queries and collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias, suggesting reliability is shaped by the generative paradigm together with training data.

Abstract

Diffusion-based Large Vision-Language Models (dLVLMs) have recently emerged as a compelling alternative to autoregressive (AR) LVLMs, offering advantages in parallel decoding, bidirectional context, and controllable generation. Despite rapid progress, their reliability properties remain largely uncharacterized. We present the first systematic reliability evaluation of hallucination and bias in dLVLMs, benchmarking six diffusion models against competitive AR baselines across four dimensions. Our key findings are: (1) dLVLMs reverse the yes-bias of AR models in binary visual queries; (2) they achieve competitive hallucination rates yet exhibit degraded linguistic quality; (3) they collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias; and (4) they exhibit accuracy collapse in multiple-choice settings when the correct option is shorter than its distractors, associated with a length prior that emerges at the first denoising step. Tokens committed at late denoising steps with low confidence further correlate with hallucinated content, pointing to a mechanistic signal unique to diffusion generation. These patterns vary across model families, suggesting reliability is shaped by the generative paradigm together with training data.

View source

Similar papers

#machine learning Preprint Sep 2026

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the tempor...

S. More, Tanuja S. Pawar · 0 citations
#artificial intelligence Preprint Aug 2026

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

ReWEIGH is a training-free decoding intervention that aggregates vocabulary ranks across visual positions and compares each candidate with a token-specific reference estimated from unlabeled images and applies a bounded penalty only to candidates that fall below their reference.

Jihae Jeong, Jun-Ha Choi, Hwanjo Yu · 0 citations
#artificial intelligence Preprint Sep 2026

Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs

Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsu...

Ya-Dong Wang, Si-Ping Yue, Yu Tian et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation

EviAnchor is proposed, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation in large vision-language models and demonstrates consistent improvements in visual grounding.

Sihang Jia, Shuliang Liu, Song-Bo Yang et al. · 0 citations
Preprint Sep 2026

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or post-hoc calibration, but face limitations in dynamic visual bias mitigation. These include inability to capture real-time...

Zhi-Peng Zhao, Zhao Wei, Peishun Liu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.