Skip to content
Open access

Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering

Jul 2026 · Multimodal Technologies and Interaction · 0 citations · 59 references

TL;DR

A unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism is proposed, which significantly improves both in-distribution accuracy and out-of-distribution robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.

Abstract

Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and generalization. Existing debiasing techniques mainly depend on post hoc evaluation or architectural modifications, while recent prompt-learning-based methods reveal new opportunities for aligning downstream tasks with pretrained models. In this work, we propose a unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism. The generative prompt component reformulates VQA as a cloze-style masked prediction problem, leveraging pretrained language priors to improve semantic grounding. Meanwhile, the fuzzing-based module actively constructs unexpected test samples during training and employs a reflection mechanism to correct biased predictions in-loop, yielding inference-time robustness without additional test-time components. Extensive experiments on VQA-v2, VQA-CP, VQA-CE, GQA-OOD, and VQA-VS demonstrate that the proposed framework significantly improves both in-distribution (ID) accuracy and out-of-distribution (OOD) robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.

Read PDF

Similar papers

Conference Jul 2026

Improving Trustworthiness in Visual Question Answering Via Question-Conditioned Cross-Modal Verification

Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this l...

Prakhar Shukla, Ankit Kumar, Pulkit Singh et al. · 0 citations
Jul 2026

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

Cross-Modal Visual Feedback (CMVF) incorporates a failure-conditioned visual diagnosis stage, in which a stronger optimizer VLM inspects each failed image without access to predictions or labels, and an error-aware aggregation stage that compresses these observations into reusable, task-level visual blind-spot patterns...

Haoyue Liu, Xiao-Yu Ma, Yeheng Chen et al. · 1 citation
Preprint Aug 2026

Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning

VeriOCRBench, a 1,800-sample human-verified benchmark built from source images drawn from 8 OCR-related datasets and spanning 8 real-world image domains, with controlled, image-grounded diagnostic tasks, is introduced, exposing a critical reliability gap in current OCR reasoning systems.

Yue Zhou, Yuan Wu, Yi Chang · 0 citations
Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Review Aug 2026

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

SABRE is established as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark, and the results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Zi-Xuan Lan, Luzhe Sun, Matthew R. Walter et al. · 0 citations
Preprint Aug 2026

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.

Fang-Zhou Chen, Shiji Zhao, Mengyan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.