Skip to content

SDGBiasBench: Benchmarking and Mitigating Vision-Language Models' Biases in Sustainable Development Goals

May 2026 · arXiv.org · Vol abs/2605.21919 · 0 citations · 47 references
Computer Science

TL;DR

This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.

Abstract

Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases. Real-world SDG monitoring further spans both qualitative judgments and quantitative estimation. However, existing benchmarks typically evaluate these aspects in isolation, obscuring systematic biases that emerge when models substitute priors for evidence. To address this gap, we propose SDGBiasBench, a large-scale benchmark suite for SDG-oriented vision-language reasoning. Spanning 500k expert-involved multiple-choice questions and 50k regression tasks, the benchmark enables comprehensive assessment of both decision-level and estimation-level bias in Vision--Language Models (VLMs). Evaluations on SDGBiasBench reveal an intrinsic SDG bias in current VLMs, where predictions are frequently driven by SDG specific priors rather than reliable multi-modal cues. To mitigate such bias, we propose CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors. CADE yields significant gains on the proposed benchmark, improving multiple-choice accuracy by up to 25% and reducing regression MAE by up to 12 points across multiple VLMs. We hope our work can foster the development of more fair and reliable AI systems for sustainable development.

View source

Similar papers

Review Aug 2026

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

SABRE is established as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark, and the results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Zi-Xuan Lan, Luzhe Sun, Matthew R. Walter et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs

This work proposes MMPCBench, a comprehensive framework for evaluating MLLMs' proactive critique competence, which adopts a hierarchical evaluation protocol to measure models' error detection, diagnosis and resolution performance, and applies alignment-aware metrics to assess the coherence between internal reasoning and final responses.

Jin-Zhe Li, Geng-Xu Li, Jinnan Li et al. · 0 citations
Preprint Aug 2026

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.

Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee et al. · 0 citations
Jul 2026

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

This is the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases).

Ioannis Sarridis, Y. Kompatsiaris, Symeon Papadopoulos · 0 citations
Preprint Aug 2026

BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning

Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs'robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation), a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics -- Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) -- to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.

S. Rehman, Muhammad Shafique · 0 citations
#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.