Skip to content
Open access

Performance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study

Sep 2026 · Applied intelligence (Boston) · Vol 56 · 0 citations · 80 references

TL;DR

A reproducible, end-to-end methodology that jointly optimizes performance and interpretability, offering practical guidance for selecting and deploying explainable deep learning models in clinical pneumonia screening is contributed.

Abstract

Chest X-ray imaging remains the most widely used radiological modality for pneumonia screening. While recent advances in deep learning have demonstrated strong diagnostic performance, the deployment of such models in real-world settings requires not only high accuracy but also robustness and interpretability across different model designs and populations. In this work, we present a comprehensive benchmarking study of multiple deep learning architectures for pneumonia detection. A distinctive methodological feature of this study is the use of two demographically distinct datasets: a pediatric chest X-ray dataset for training, and an independent adult population dataset for external validation. All models were evaluated under a standardized cross-validated protocol. Beyond predictive metrics, we conduct an extensive eXplainable Artificial Intelligence (XAI) analysis assessing both qualitative and quantitative properties, including explanation stability and localization fidelity. Results show that model design choices significantly influence both predictive performance and explanation behavior. In particular, certain architectures consistently achieved superior performance while producing more focused and stable explanations. The performance-interpretability ranking is preserved under external validation on a demographically distinct adult cohort, providing evidence that performance-interpretability coupling is robust to population shift and supports the generalizability of the proposed framework beyond the pediatric training distribution. This work contributes a reproducible, end-to-end methodology that jointly optimizes performance and interpretability, offering practical guidance for selecting and deploying explainable deep learning models in clinical pneumonia screening. All code and data are publicly available at: https://github.com/SynergIA-Lab/pneumoniacnn

Read PDF

Similar papers

Open access Sep 2026

INFORMER-Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies.

Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only...

Shelley Zixin Shu, Aurélie Pahud de Mortanges, A. Poellinger et al. · 0 citations
Open access Aug 2026

Analysis of CNN-Based Deep Learning Architectures and Performance Enhancement Strategies for Pneumonia Classification Using Chest X-Ray Images

Background: Deep learning models, particularly convolutional neural networks (CNNs), have shown promising performance for pneumonia detection using chest X-ray images. However, the impact of preprocessing, architecture selection, data augmentation, and ensemble strategies has not been systematically evaluated. This stu...

YongJun Kim, Ji-Yeoun Lee · 0 citations
Open access Aug 2026

Reducing False-Negative Risk in Automated Pneumonia Screening Using Attention-Guided Deep Learning

An attention-enhanced deep learning framework for clinically accurate pneumonia identification from chest imaging radiology that combines a self-attention mechanism with a pretrained VGG16 backbone is proposed and tested against many cutting-edge convolutional neural network architectures.

Mohini Gahlot, Pinaki Ghosh · 0 citations
Conference Aug 2026

Causally Constrained and Anatomically Grounded Interpretability in Deep Learning for Chest Radiograph Diagnosis

Although deep learning systems have proved to have high potential in classifying chest radiographs, their low level of transparency still limits their widespread clinical use on a regular basis. Most current models produce visual explanations that are generated post-prediction, and frequently point out regions that do...

U. Venkat Subbaiah, T. Narayanan, B. Kumar · 0 citations
#artificial intelligence Preprint Sep 2026

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as...

B. Renoust, P. Baudot, T. Foriel et al. · 0 citations
Conference Aug 2026

Benchmarking Test-Time Adaptation for Multi-Label Chest X-ray Classification under Distribution Shift

Although AI models achieve impressive performance on chest X-ray benchmarks, their deployment in real-world clinical settings remains challenging due to performance degradation under unpredictable conditions. To understand how these AI models can adapt in practice, we systematically evaluate several Test-Time Adaptatio...

Duc-Ngo Van, Kim-Hung Le · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.