Skip to content
Open access

Hybrid MobileNetV2–vision transformer approach for multi-label classification of chest X-rays

Aug 2026 · Scientific Reports · 0 citations

TL;DR

A hybrid MobileNetV2–Vision Transformer (ViT) framework for multi-label classification of CXR images into 14 disease categories on NIH CXR14 dataset is introduced, which adaptively optimizes key hyperparameters of the MobileNetV2–ViT framework to achieve improved accuracy, faster convergence, and enhanced computational efficiency compared to conventional approaches.

Abstract

The precise identification of thoracic disorders using chest X-rays (CXRs) is essential for favorable clinical diagnosis. However, it continues to be difficult because of inter-class similarities and overlapping symptoms. This study introduces a hybrid MobileNetV2–Vision Transformer (ViT) framework for multi-label classification of CXR images into 14 disease categories on NIH CXR14 dataset. MobileNetV2 serves as a lightweight feature extractor to obtain distinctive spatial representations that are tokenized and enhanced with positional encodings. The encoded features were processed via transformer layers, wherein self-attention mechanisms captured long-range interdependence and global contextual relationships. An output layer based on the sigmoid function facilitated the concurrent prediction of various illness labels. The novelty of the proposed work lies in the PSO-tuned hybrid architecture, which adaptively optimizes key hyperparameters of the MobileNetV2–ViT framework to achieve improved accuracy, faster convergence, and enhanced computational efficiency compared to conventional approaches. This implementation uses an economical, GPU-accelerated instance tailored for inference workloads. It is equipped with a single NVIDIA A10G GPU with 24 GB of dedicated GPU memory, and the proposed work is carried out on AWS SageMaker. The experimental assessment employing five-fold cross-validation showed strong performance, with an average accuracy of 94.6%, precision of 88.1%, recall of 85.9%, specificity of 94.17%, and F1-score of 91.1%. Grad-CAM representations increased interpretability by identifying disease-relevant areas. The results underscore the efficacy of the proposed hybrid technique as a dependable instrument for computer-aided diagnosis of thoracic illnesses.

Read PDF

Similar papers

Open access Aug 2026

XAI-Enhanced Hybrid CNN–Transformer Framework For Multi-Class Lung Disease Classification

The results suggest that the hybrid CNN-Transformer model provides a strong level of diagnostic accuracy and meaningfully understood visual rationale so it can serve as an excellent decision support mechanism for hospitals and radiologists in their daily operations.

Prasanna Pabba, N. S. Chaitanya, M. Ravikanth et al. · 0 citations
Conference Open access 2026

Vision Transformer–Driven Feature Extraction for Explainable Pneumonia and COVID-19 Detection from Chest X-Ray Images

In order to diagnose respiratory disorders such as pneumonia and COVID-19, X-ray imaging of the chest is essential. On the other hand, radiologists could differ in their approaches and the amount of time it takes to manually analyze radiographic pictures. Recent innovations in deep learning have greatly enhanced automa...

P. V. Naga Lakshmi, K. Vedavathi · 0 citations
Conference Aug 2026

Multi-Scale DenseNet with Hybrid Shuffle-Spatial Attention for Robust Lung Disease Classification on Chest X-Rays

Lung diseases such as pneumonia, tuberculosis, and the current coronavirus (COVID-19) are significant causes of morbidity and mortality in the world. Early diagnosis and accurate diagnosis based on chest X-rays (CXR) is very important for effective treatment, but manual interpretation takes time and can be prone to err...

V. Nandhini, R. Parameswari · 0 citations
Conference Aug 2026

An Explainable CBAM Enhanced DenseNet121 Framework for Multi-Class Lung Cancer Classification Using CT Scans

Due to its late identification and challenging diagnosis, lung cancer continues to be one of the top causes of death for cancer patients globally, positioning it as one of the most critical concerns. Timely identification of cancerous nodules is essential for enhancing the patient’s survival likelihood CT image analysi...

S. Jegadeesan, S. Matheswaran, R. Palanivelrajan · 0 citations
Open access Aug 2026

A monotonic multi-expert Vision Transformer for clinically reliable chest X-ray classification

Introduction Chest X-ray (CXR)-based recognition of pulmonary diseases remains challenging due to overlapping radiographic patterns, class imbalance, and variability across imaging sources. These factors often lead to unstable performance in large-scale multi-class classification tasks. Methods In this study, pulmonary...

Xiang Wu, Yogesh H. Bhosale, H. Muhammad et al. · 0 citations
Open access Sep 2026

LPC-transformer: a training strategy optimized framework for multiclass pneumonia medical image classification

Lung imaging enables direct visualization of lesions and is vital for pneumonia diagnosis. However, current deep learning models show unsatisfactory classification accuracy and generalization for multi-class lung disease recognition, limiting their clinical decision-support value. To tackle this issue, we present LPC-T...

Hong-Wei Chang, Tao Hu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.