Skip to content
Conference

Cross-Model Attention-based Vision–Language Fusion for Interpretable Multi-Label Chest Disease Diagnosis

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1481-1486 · 0 citations · 17 references

Abstract

Chest diseases including pneumonia, cardiomegaly, pleural effusion, atelectasis and consolidation are major public health problems around the world that need to be diagnosed quickly and accurately to provide appropriate treatment. In this paper, we propose a cross-modal attention-based multimodal deep learning method that integrates both chest X-ray (CXR) image and related radiology report for automated multi-label classification of chest diseases. A convolutional network is used as the backbone for the visual representations while semantic textual information is extracted using the BioBERT-based language modeling. To enable meaningful interaction between image and text features and enhance feature integration and diagnostic inference, a cross-modal attention module is introduced. For better explanation of predictions, both visual by Grad-CAM and textual by SHAP are used. Extensive experiments conducted on MIMIC-CXR and NIH ChestX-Ray14 datasets show that the proposed framework outperforms other unimodal and conventional fusion methods by not only accuracy but also AUC, recall, and overall robustness. The results show that multimodal vision–language learning is a clinically applicable, interpretable, and effective approach to computer-aided diagnosis of chest diseases.

View source

Similar papers

Review Open access Sep 2026

A Unified Explainable AI Framework for Multimodal Chest X-Ray and Clinical Note Diagnosis

Chest radiography is the most common first-line imaging modality for diagnosing respiratory disease, but deep learning classifiers built for this task are largely opaque, which restricts their clinical adoption. This paper proposes a multimodal explainable-AI (XAI) framework that couples a DenseNet-121 convolutional ne...

Shobhanjaly P. Nair, Lavanya Jayaraman · 0 citations
Open access Oct 2026

An attention-based multimodal hybrid deep learning framework for pneumonia classification using chest X-rays and radiologist-annotated severity labels

Adequate classification of pneumonia is essential for timely disease management and effective resource allocation in healthcare systems. As pneumonia presents with varying severity and its assessment utilizing chest radiographs is challenging, specifically detecting MILD and NORMAL-PCR+ cases, where chest radiographs m...

Khalida Anum, Muhammad Nouman Noor, T. Akram et al. · 0 citations
Conference Open access 2026

Vision Transformer–Driven Feature Extraction for Explainable Pneumonia and COVID-19 Detection from Chest X-Ray Images

A Vision Transformer (ViT)-based feature extraction system embedded in the GenMAT-Net framework to classify X-ray images of the chest automatically and provide a promising solution to intelligent computer-aided diagnostic systems in the medical imaging field is suggested.

P. V. Naga Lakshmi, K. Vedavathi · 0 citations
Open access Aug 2026

XAI-Enhanced Hybrid CNN–Transformer Framework For Multi-Class Lung Disease Classification

The results suggest that the hybrid CNN-Transformer model provides a strong level of diagnostic accuracy and meaningfully understood visual rationale so it can serve as an excellent decision support mechanism for hospitals and radiologists in their daily operations.

Prasanna Pabba, N. S. Chaitanya, M. Ravikanth et al. · 0 citations
Open access Aug 2026

Hybrid MobileNetV2–vision transformer approach for multi-label classification of chest X-rays

A hybrid MobileNetV2–Vision Transformer (ViT) framework for multi-label classification of CXR images into 14 disease categories on NIH CXR14 dataset is introduced, which adaptively optimizes key hyperparameters of the MobileNetV2–ViT framework to achieve improved accuracy, faster convergence, and enhanced computational...

R. Raj, Pavan M. P. Kumar, K. N. Manjunath et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.