Skip to content
Open access

An Automated, Contamination-Controlled VQA Benchmark for Evaluating Vision-Language Models on 3D Oncology Imaging

Sep 2026 · Research Square · 0 citations · 18 references
Medicine

Abstract

Abstract Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that builds multiple-choice benchmarks directly from paired private radiology reports and three-dimensional oncology imaging. It generates two complementary question types: schema-driven items populated deterministically from established reporting frameworks, and report-derived items verified against the source text. Because the source reports are private and single-institution, the resulting questions cannot have entered any model’s pretraining data, controlling instance-level contamination by construction. We applied the pipeline to four oncology cohorts (liver CT, liver MRI, lung CT and brain MRI; 2,509 cases and 33,870 questions) and evaluated five contemporary VLMs zero-shot. Accuracies ranged from 0.28 to 0.81 and no model was reliable across cohorts. Replacing the image with a blank input changed accuracy little for the highest-scoring models, and on schema-driven brain questions frontier models scored 0.17–0.19 higher on public than on matched private imaging. The benchmark thus separates image-dependent performance from question-text and dataset-familiarity effects, and the released pipeline lets institutions regenerate it on their own reports and images.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.