ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.
Abstract
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.
Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that builds multip...
Bo Liu, Han Gu, Xiang-Rui Li et al.· Research Square· 0 citations
Across six medical multimodal MCQ datasets, this work separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key, showing that medical image-reasoning claims require route-level evidence.
Ben Wang, Yi-Fan Zhang, Jia-Qing Yu et al.· 0 citations
In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth a...
Jian-Yu Sun, Zhen-Xuan Zhang, Guang Yang et al.· 0 citations
It is suggested that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.
Hermione Warr, Harry Anthony, Lilli J. Freischem et al.· 0 citations
Introduction: Multimodal large language models, including GPT-4 Omni (GPT-4o), have been applied for facilitating the healthcare process, but their capacity to interpret thyroid sonography images to aid report generation, as well as ways for improvements, are unclear. Methods: 120 thyroid nodules were retrospectively i...
Ze-Bang Yang, Tong-Yi Huang, Lei Huang et al.· Current medical imaging· 1 citation
We developed a national questionnaire to assess radiology reports (RR) physicians’ reading habits and identify their expectations towards RR, focusing on CT and MRI reports. An anonymized 20-item questionnaire was nationally distributed in France via practice and hospital mailing lists and one social media group, and i...
Yasmine Kassab, Matthieu Bailly, Sixtine Brabant et al.· Insights into Imaging· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.