Skip to content
Open access

Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis

Sep 2026 · Journal of Artificial Intelligence in Bioinformatics · Vol 2, pp. 47-54 · 0 citations · 24 references

TL;DR

A fully supervised framework that couples a Vision Transformer encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset indicates that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.

Abstract

Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformerbackbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.

Read PDF

Similar papers

Open access 2026

A Hybrid Vision Transformer and U-Net Framework for Automated Cardiac MRI Segmentation and Abnormality Detection

Cardiovascular diseases remain a leading cause of mortality worldwide, and accurate segmentation of cardiac structures from MRI is critical for clinical diagnosis. We propose a hybrid framework that integrates a Vision Transformer encoder with a U-Net decoder for automated cardiac MRI segmentation and abnormality detec...

Prachi Khune · 0 citations
Open access Sep 2026

VentrEX: An Anatomically Guided Deep Learning Pipeline for Ventricular Segmentation in Cine Cardiac MRI

Automated segmentation of the left and right ventricles (LVs and RVs) in cine cardiac MRI (CMR) underpins reliable volumetry and mass estimation. However, papillary muscles and trabeculae (PM/T) introduce clinically meaningful variability and exacerbate cross-dataset domain shift. We present VentrEX, an anatomically gu...

Abla Bedoui, Julieta Anahí Rancati, Ignacio Lugones et al. · 0 citations
Open access Aug 2026

MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation

To overcome the architectural mismatch inherent in existing hybrid networks, a novel Scaling Feature Pyramid (SFP) is proposed, which effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring th...

Zhi-Yu Ye, Hai-Rong Zheng, Tong Zhang · 0 citations
Conference Open access 2026

CDSS for Automated Cardiac MRI Diagnosis Using an Explainable Ensemble Deep Learning Model

It is demonstrated that architectural diversity combined with probabilistic aggregation constitutes an effective and interpretable strategy for reliable cardiac MRI diagnosis in clinical decision support systems.

Soukaina Ait Ouaoures, Hayat Bihri, Salma Azzouzi et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI

MR-JEPA is presented, a self-supervised video foundation model for CMR that extends LeJEPA to 3D spatiotemporal inputs through tubelet tokenization, spatiotemporal masking augmentation, and initialization from a 2D CMR foundation model.

Athira J. Jacob, Puneet Sharma, D. Comaniciu et al. · 0 citations
Sep 2026

CardioFusion-XAI: a robust multi-scale and explainable framework for cardiac MRI-based disease classification

CardioFusion-XAI provides accurate and interpretable cardiac MRI classification and WaveCLAHE-Net integrates wavelet-based enhancement and contrast-limited adaptive histogram equalisation to improve image quality.

Shwetambari Borade, Saraswati Mishra, R. Vairagade et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.