A Clinically Validated Explainable Federated Multi-Modal Learning Framework for Privacy Preserving Cardiac Arrhythmia Classification in Heterogeneous Data Environments
Abstract
Cardiac arrhythmia is a leading cause of morbidity worldwide, and automated electrocardiogram (ECG) classification has become an essential tool for early screening and triage. Two barriers separate research prototypes from clinical deployment: patient data cannot be centralized due to privacy regulations (HIPAA, GDPR), and clinicians do not trust opaque “black-box” predictions. We propose FedMM-XAI, a federated multi-modal learning framework that fuses raw ECG waveforms with structured electronic health record (EHR) features through a cross-modal attention mechanism, trains collaboratively across heterogeneous (non-IID) client sites without sharing raw data, and produces per-prediction explanations via Grad-CAM on the ECG stream and SHAP on the tabular stream. Differential privacy (DP-SGD) and secure aggregation are integrated at the communication layer to provide formal privacy guarantees. Evaluated on a heterogeneous, ten-client partitioning of the MIT-BIH Arrhythmia Database and PTB-XL with Dirichlet-distributed non-IID label skew $(\alpha \in\{0.1,0.5,1.0,5.0\})$, FedMM-XAI achieves 97.42% accuracy and a macro-F1 of 96.97%, outperforming single-modality federated baselines (FedAvg: 93.15%, FedProx: 94.02%) while remaining within 0.7 percentage points of a centralized, non-federated upper bound. A structured clinician-review pilot, in which board-style cardiology reviewers rated explanation relevance on a subset of predictions, indicated substantially higher perceived trust for the explainable variant than for a non-explainable baseline. Under a per-round privacy budget of $\varepsilon=2$, accuracy degrades by only 1.1-3.1% depending on data heterogeneity. These results suggest that explainable, privacy-preserving federated multi-modal learning is a practical path toward clinically deployable arrhythmia screening tools.