An Automated, Contamination-Controlled VQA Benchmark for Evaluating Vision-Language Models on 3D Oncology Imaging
Abstract Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that buil...