An Attention-Based Multi-Modal Network Framework for Cardiovascular Disease Detection and Grading
Cardiovascular Disease (CVD) is the most common cause of mortality worldwide, so reliable tools are needed to accurately diagnose it to provide timely clinical interventions. Traditional forms of diagnosis relied solely on single-modality data, used basic fusion approaches at diagnosis, or did not consider how to complement data across heterogeneous modalities. This study presents a novel Hierarchical Attention-based Representation learning with Multi-modal Network (HARM-Net) framework to grade the severity of CVD. The proposed framework combines medical imaging, physiological signals, electronic health records, and demographic data. It consists of a five-step process that includes using modality-specific encoders, self-supervised pre-training via multi-modal contrastive learning, hierarchical cross-attention fusion, compressing deep features, and a modality-specific adaptive stacking ensemble classification model. Extensive experimental results on the MultiD4CAD dataset demonstrated that the HARM-Net framework outperformed other models with an accuracy of 93.2%, an F1-score of 91.7%, and a ROC-AUC of 0.967. The results of an ablation study demonstrate the significance of each component and the specific contribution of cross-attention fusion to improved model performance for accurate, multimodal diagnosis of CVD.