Skip to content

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

Jul 2026 · arXiv.org · Vol abs/2607.26333 · 0 citations
Engineering Computer Science

TL;DR

This work systematically investigates how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA), and shows that for supervised image classifiers, changing the label source leads to substantial differences not only in performance estimates but also in model rankings.

Abstract

Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.

View source

Similar papers

Open access Aug 2026

A monotonic multi-expert Vision Transformer for clinically reliable chest X-ray classification

Overall, the proposed QMIX-ViT framework provides a structured approach for multi-class chest X-ray classification under heterogeneous data conditions and suggests that disease-group expert decomposition and monotonic fusion can reduce inter-class interference and improve decision-level stability.

Xiang Wu, Yogesh H. Bhosale, H. Muhammad et al. · 0 citations
Open access Sep 2026

Automated chest X-ray disease screening using large language models and deep convolutional neural networks on the MIMIC-CXR dataset

It is demonstrated that LLMs can be effectively employed to generate supervision labels for medical imaging tasks and that the proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability.

Qing-Yuan Zhang, Pardeep Vasudev, Kezhi Li et al. · 0 citations
Sep 2026

SpineXR-VQA: A Clinically-Validated Visual Question Answering Dataset for Spine X-Rays

This resource paper introduces SpineXR-VQA, an open-source, clinically grounded, and verified benchmark comprising 2,187 X-rays and 8,272 expert-verified, open-ended Question-Answer pairs that benchmark 15 state-of-the-art Multimodal Large Language Models (MLLMs), including proprietary systems such as Claude Sonnet, GP...

Deepali Mishra, Dr.Sorayouth Chumnanvej, Vikas Trivedi et al. · 0 citations
Review Open access Aug 2026

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings...

S. Alhebshi, H. Khalifa, T. D. Pham · 0 citations
Open access Sep 2026

Performance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study

A reproducible, end-to-end methodology that jointly optimizes performance and interpretability, offering practical guidance for selecting and deploying explainable deep learning models in clinical pneumonia screening is contributed.

Francisco A. Gómez-Vela, Aurelio López-Fernández, F. Divina et al. · 0 citations
Conference Jul 2026

Optimizing InceptionV3 through Network Pruning for Lung Disease Detection in Chest X-Ray Images

Lung disease remains a major global health concern, and accurate diagnosis using chest X-ray images plays a crucial role in supporting effective clinical decision-making. The contribution of this work lies in empirically demonstrating how internal redundancy removal through standard magnitude-based pruning can improve...

Joshua Pinem, Widi Astuti, A. Adiwijaya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.