Aug 2026· Diagnostics· Vol 16· 0 citations· 43 references
Medicine
TL;DR
A compact and interpretable framework for five-superclass multi-label ECG diagnosis that integrates leakage-aware development, threshold-controlled testing, frozen external validation, and multimethod interpretability is developed.
Abstract
Background: Automated interpretation of 12-lead electrocardiograms (ECGs) remains challenging because multiple abnormalities may coexist and appear in selected leads or brief waveform segments. We developed a compact and interpretable framework for five-superclass multi-label ECG diagnosis. Methods: We evaluated PTB-XL records using the official fold protocol, with folds 1–8 for training, fold 9 for validation monitoring and class-specific threshold selection, and fold 10 for independent internal testing. We then evaluated the frozen 1.33-million-parameter InceptionTime–CNN–BiGRU–Transformer model and validation-derived thresholds on 15,931 Ningbo ECGs without retraining, recalibration, or external threshold adjustment. We also examined calibration, demographic subgroups, computational efficiency, and complementary ECG-domain attribution methods. Results: Macro-AUROC reached 90.83% on PTB-XL fold 10 and 88.85% on Ningbo, indicating generally consistent diagnostic ranking with a modest reduction during external evaluation. Sensitivity analysis showed that CNN-only outperformed the frozen primary model on five of six endpoints. Attribution analyses highlighted qualitatively plausible lead and temporal patterns in representative examples, while calibration and subgroup analyses further characterized model behavior under dataset shift. Conclusions: Our framework integrates leakage-aware development, threshold-controlled testing, frozen external validation, and multimethod interpretability. These findings support its further prospective, locally calibrated evaluation as a potential aid for multi-label ECG interpretation.
X- Beat is presented, an explainable and reliability-aware benchmark framework for ECG image classification designed to support trustworthy AI systems in healthcare and provides a structured and reproducible bench- mark for evaluating both predictive performance and explanation reliability in ECG image classification.
Mohammad Sadman Tahsin, Haitham Y. Adarbah, A. Noore· 0 citations
An end-to-end Multi-Scale Hybrid Transformer model was developed to capture both local morphological features and long-range temporal dependencies across multiple leads to demonstrate robust performance in detecting Brugada patterns from raw ECG inputs.
Jun-Mo An, Zi-Yu Li, D. Dzikowicz et al.· 1 citation
It is demonstrated that hybrid spatio-temporal architectures can achieve diagnostic performance comparable to strong convo-lutional baselines while offering significantly improved trans-parency through a quantitative comparison of representative models and visual analysis of performance—interpretability trade-offs.
Rhivu Dutta, N. Suma· JIMS8I - International Journ...· 0 citations
ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.
M. Hassannia, Matthew A. Reyna, R. Sameni· 0 citations
Background Electrocardiography (ECG) is the most widely deployed cardiac screening modality, but demographic, acquisition and label-space heterogeneity across cohorts obstructs clinical translation. Cross-population studies rarely quantify which conditions transfer, how much target labelling is needed, or whether the a...
Yan Wang, Yang Lu· Frontiers in Cardiovascular...· 0 citations
This ECG-based deep learning model demonstrates good discrimination and calibration across AMI subtypes, indicating potential clinical utility for rapid risk stratification and early cardiology intervention.
Tobias Zimmermann, Ivo Strebel, P. López-Ayala et al.· npj Digital Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.