Aug 2026· Academic Radiology· 0 citations· 15 references
Medicine
TL;DR
Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained convolutional neural network remains robust within its training domain.
Abstract
Rationale
AND
Objectives
Multimodal large language models (LLMs) are increasingly applied to image-based radiology tasks, but their diagnostic accuracy across clinically distinct populations remains poorly characterized. We quantified age-related differences in zero-shot LLM performance for pneumonia detection on pediatric vs. adult chest radiographs and compared generalization gaps with a domain-trained convolutional neural network (CNN) baseline.
Materials And Methods
GPT-5.2 (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Pro (Google) were evaluated zero-shot on balanced pediatric and adult test sets of frontal chest radiographs (n = 1000 each; 500 pneumonia/500 normal). Cohort-specific InceptionV3 CNNs were trained on the remaining development-pool images (pediatric n = 4715; adult n = 13,863) and evaluated on the same test sets. Performance was assessed using the Matthews correlation coefficient (MCC) with 95% bootstrap confidence intervals (CIs); domain shift was quantified as Δ = Adult - Pediatric.
Results
In pediatrics, the CNN outperformed all LLMs (MCC 0.799, 95% CI 0.766-0.832) vs. GPT-5.2 (0.484, 0.436-0.532), Claude Opus 4.5 (0.470, 0.418-0.521), and Gemini 2.5 Pro (0.272, 0.224-0.316). In adults, all models improved, but the CNN remained best (MCC 0.850, 0.816-0.882). Age-related gains were larger for LLMs (ΔMCC +0.220 [95% CI 0.160-0.281] to +0.466 [0.407-0.525]) than for the CNN (ΔMCC +0.051 [0.005-0.098]), driven mainly by specificity increases.
Conclusion
Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.
Large-scale adult chest radiograph datasets are commonly used to develop deep learning models for pneumonia detection, whereas pediatric pneumonia datasets remain smaller, more heterogeneous, and less widely available. Because pediatric chest radiographs differ from adult radiographs in anatomical proportions, diseas...
Yu-Tao Li, Junghun Kim, Sang-Il Choi· Scientific Reports· 0 citations
Background/Objectives: Pneumonia remains a leading cause of childhood morbidity and mortality worldwide. Accurate interpretation of pediatric chest radiographs is challenging because of anatomical variability, subtle radiographic findings, and inter-observer variability. This study evaluates different CNN–Transformer e...
Ece Meltem Yalçın, Hayriye Tanyıldız, Serpil Aslan et al.· Diagnostics· 1 citation
It is demonstrated that LLMs can be effectively employed to generate supervision labels for medical imaging tasks and that the proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability.
Qing-Yuan Zhang, Pardeep Vasudev, Kezhi Li et al.· Frontiers in Digital Health· 0 citations
This study aims to compare the classification performance and computational complexity of a CNN and a pre-trained ResNet50 model using transfer learning for binary pneumonia classification on chest X-ray images and shows that the CNN outperforms the ResNet50 across all classification metrics.
Tam Pran Noto Noto, Supatman Supatman· Jurnal Riset Informatika· 0 citations
This study examines a deep learning technique called a Convolutional Neural Network that looks at chest X-rays to identify pneumonia and suggests that CNN-based models have the capacity to aid radiologists in the early diagnosis to ease the medical intervention and minimize diagnostic errors.
M. Devi, Tanya, Aradhya Mittal et al.· Proceedings of the 1st Inter...· 0 citations
This study investigates whether class imbalance can be mitigated through random under-sampling, random over-sampling, and medically informed image augmentation when training three representative deep learning architectures: EfficientNet, Vision Transformer, and Swin Transformer.
Pablo Ormeño-Arriagada, Valentina Zúñiga, Carlos A. Toro et al.· Cancers· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.