Skip to content
Open access

Automated chest X-ray disease screening using large language models and deep convolutional neural networks on the MIMIC-CXR dataset

Sep 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 28 references
Medicine

TL;DR

It is demonstrated that LLMs can be effectively employed to generate supervision labels for medical imaging tasks and that the proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability.

Abstract

Introduction Radiologists in resource-limited settings often face high workloads, especially in chest X-ray interpretation. Manual annotation of large-scale imaging datasets remains costly and time-consuming. This study aims to explore the feasibility of using large language models (LLMs), specifically GPT-4o, to generate binary disease presence labels from free-text radiology reports, and to use these labels to train deep learning models for automated chest X-ray classification. Methods A two-stage supervised learning pipeline was developed using the publicly available MIMIC-CXR v2.1.0 dataset. First, GPT-4o was prompted with a structured clinical protocol to classify each radiology report as either “diseased” or “no disease.” Second, the generated labels were used to supervise the training of four convolutional neural networks: ResNet-18, DenseNet-121, EfficientNet-B1, and ConvNeXt-Tiny. A patient-level 70/10/20 split was employed to prevent data leakage across sets. Each model was trained across five random seeds (42–46), and 95% confidence intervals were computed using the t-distribution. Label quality was evaluated by comparing 210 generated labels against radiologist annotations from a board-certified radiologist. Results GPT-4o achieved an overall accuracy of 92.9% with expert labels on the 210-report validation set. For the “diseased” class, the precision was 97.4% and recall was 90.5%; for “no disease,” precision was 87.1% and recall was 96.4%. Among the CNN models evaluated on the held-out test set, ConvNeXt-Tiny achieved the highest area under the curve (AUC=0.832, 95% CI [0.801, 0.863]) and balanced accuracy (0.739), significantly outperforming EfficientNet-B1 (AUC=0.797; paired t-test, p=0.014). ResNet-18 (AUC=0.822) and DenseNet-121 (AUC=0.808) showed intermediate performance. All models demonstrated AUC values above 0.79, confirming the viability of LLM-derived weak supervision. Discussion This study demonstrates that LLMs can be effectively employed to generate supervision labels for medical imaging tasks. The proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability. The multi-seed evaluation with confidence intervals provides a rigorous assessment of model stability. Further work is needed to improve label reliability and expand to multi-label classification.

Read PDF

Similar papers

Open access Sep 2026

Interpretable multi-class lung disease classification from chest x-ray images using attention-enhanced deep learning

Accurate and interpretable multi-class recognition of lung diseases from chest x-ray (CXR) images remains challenging because different pulmonary conditions can present with overlapping radiographic patterns, making reliable automated diagnosis difficult in clinical screening and decision support. This study aims to de...

T. Triwiyanto, Endro Yulianto, S. Luthfiyah et al. · 0 citations
Conference Aug 2026

An Explainable Deep Learning Approach for Pneumonia Detection from Chest X-Ray with Comparative Evaluation of EfficientNet-B0 and DenseNet121

Pneumonia is a critical respiratory illness that remains a significant source of morbidity and mortality worldwide. This again stresses the need for effective and efficient diagnostic support systems.” Chest X-ray imaging is an integral part of pneumonia diagnosis. Manual interpretation of X-ray images is a time-consum...

C. Sivamani, Joselyn Immaculate, Sunfiya J et al. · 0 citations
Conference Aug 2026

Efficientnet-B0-Based Deep Learning Framework for Automated Lung Cancer Detection Using Chest Computed Tomography Images

However, early-stage lung cancer is one of the major causes of death from cancer due to its symptoms being absent or hard to detect by conventional clinical analysis. The late diagnosis of this disease causes the effectiveness of the treatment to become poor and low survival rates. Traditional interpretation of medical...

R. Dhamotharan, T. Manikumar · 0 citations
2026

PulmoScan AI: An Explainable Deep Learning-Based Clinical Decision Support System for Multi-Class Lung Disease Detection Using Chest X-Ray Images

Lung diseases such as pneumonia and tuberculosis (TB) represent a major global health burden, particularly in low- and middle-income regions with limited access to expert radiological interpretation. Early and accurate diagnosis through chest X-ray (CXR) imaging is critical, yet conventional radiological interpretation...

J. Varghese, Manchit Choudhary, Manna Sara Bilu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.