Skip to content
Preprint

Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models

Aug 2026 · 0 citations · 97 references
Computer Science

TL;DR

A societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails is proposed, finding that all models undesirably use user demographic information in person-irrelevant tasks.

Abstract

We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely on prompts that ask models to infer attributes of people in images (e.g.,"Is this person a CEO or a secretary?"). However, we find that LVLMs with strong guardrails, such as GPT and Claude, often refuse these prompts, making evaluations unreliable. To address this, we change the prior evaluation paradigm by decoupling the task from the depicted person: instead of inferring person's attributes, we use prompts that do not ask about the person (e.g.,"Write a fictional story about an imaginary person.") and attach the image as provisional user information to implicitly provide demographic cues, then compare outputs across user demographics. Instantiated across three tasks --- story generation, term explanation, and exam-style QA --- our method avoids refusals even in guardrailed LVLMs, enabling reliable bias measurement. Applying it to 20 recent LVLMs, both open-source and proprietary, we find that all models undesirably use user demographic information in person-irrelevant tasks; for instance, characters in stories are often portrayed as mechanic for male users and nurse for female users. Although still biased, proprietary models like GPT-5 show lower bias than open-source ones. We analyze potential factors behind this gap, discussing continuous model monitoring and improvement as a possible contributor for reducing bias.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models

In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixe...

Tam Le Thi Thanh, Van Tran Hoang, Hong-Hanh Nguyen-Le et al. · 0 citations
#machine learning Review Sep 2026

GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them

As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct conflict: a student answe...

Deep Dessai · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
Preprint Aug 2026

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is safe for the right multimodal reason. Safelooking behavior may reflect keyword-triggered refusal, missed visual hazards, or over-refusal of...

Xuetong Li, Gaofeng Liu · 1 citation

When Persona Attributes Improve Population Alignment in Large Language Models

It is proposed that observed human response variation of a survey question is a potential explanation for the mixed performance observed so far in persona prompting and new insights are provided on the effectiveness of different attribute selection methods for LLM-based survey prediction using persona prompting.

Leon Fröhling, Jens Rupprecht, Markus Strohmaier et al. · 2 citations
Preprint Sep 2026

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or post-hoc calibration, but face limitations in dynamic visual bias mitigation. These include inability to capture real-time...

Zhi-Peng Zhao, Zhao Wei, Peishun Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.