Sep 2026· Int. J. Medical Informatics· Vol 217, pp.
106504
· 0 citations· 63 references
MedicineComputer Science
TL;DR
In light of the pervasive methodological limitations identified, including high analytic risk of bias, absence of external validation, and lack of model interpretability, claims of ML superiority over CHA2DS2-VASc must be interpreted with caution.
Abstract
Background
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and confers a four to fivefold increase in ischemic stroke risk, accounting for approximately 15 - 20% of all stroke events globally. Despite this burden, the predominant risk stratification tool, the CHA2DS2-VASc score, achieves only modest discrimination, constrained by its static, additive architecture that cannot capture the nonlinear, high-dimensional interactions inherent in real-world electronic health record (EHR) data. This evidence gap creates a dual clinical hazard: under-anticoagulation in high-risk patients and unnecessary bleeding exposure in those whose risk is overestimated. This study aimed to systematically evaluate the predictive performance, methodological rigor, and clinical readiness of machine learning (ML) models derived from EHR data for the prediction of ischemic stroke in patients with AF.
Methods
A systematic search of PubMed, Embase, Scopus, and Web of Science was conducted from inception through September 2025, following PRISMA 2020 guidelines. Studies were eligible if they developed or validated ML models for ischemic stroke prediction using EHR data in adults with AF and reported at least one quantitative performance metric. Methodological quality was assessed using the PROBAST and TRIPOD-AI frameworks.
Results
Eight studies (2017 to 2024) encompassing 809,523 patients across seven countries were included. Supervised ensemble methods consistently outperformed CHA2DS2-VASc, with AUROCs ranging from 0.66 to 0.91 versus 0.54 to 0.68 for the traditional score. However, performance varied substantially: several models achieved only marginal gains (AUROC 0.63 - 0.69), and the AUROC range reflects pronounced heterogeneity rather than uniform superiority. Critical barriers persist - only one study performed external validation; fewer than half applied explainable AI techniques; class imbalance was rarely addressed; and 88% of studies received a high risk of bias rating in the analysis domain under PROBAST, a finding that substantially limits confidence in the reported performance estimates.
Conclusion
In light of the pervasive methodological limitations identified, including high analytic risk of bias, absence of external validation, and lack of model interpretability, claims of ML superiority over CHA2DS2-VASc must be interpreted with caution. While ML models demonstrate potential discriminative improvements, current evidence is insufficient to support clinical adoption. Translating algorithmic promise into bedside impact requires dynamic longitudinal modeling, rigorous multisite external validation, transparent risk attribution, and prospective evaluation within real-world EHR workflows.
In patients with cryptogenic stroke receiving an ICM, the ECG-AI score showed modest discrimination for AF detection, outperforming CHA2DS2-VA and HAVOC, but not Brown ESUS-AF, which indicates a possible role for AI-driven ECG analysis in risk stratification.
F. Wouters, M. Barthels, J. Vranken et al.· Digital Health· 0 citations
CIED‐detected AF burden is strongly associated with progression to persistent AF, and ML‐based analysis of 6‐month device data enables accurate, point‐in‐time risk stratification to support earlier and more targeted clinical management.
A. Nakonechnyi, Shaul Geliaks, I. Goldenberg et al.· Annals of Noninvasive Electr...· 0 citations
The EHR-based machine learning model, FIND-AF 2.0, identifies a high-risk subpopulation for AF diagnosis among patients at elevated risk of stroke and could enable scalable, EHR-driven, risk-guided AF screening.
R. Nadarajah, Jianhua Wu, A. Wahab et al.· Circulation· 0 citations
The guideline-endorsed CHA2DS2-VA score showed the lowest discrimination, and patients classified as intermediate-risk using this score had stroke incidence below the treatment threshold, indicating truly low-risk patients in Australia.
N. Zubrzycki, K. Hyun, Erdahl Teber et al.· Open Heart· 0 citations
This data-driven, interpretable XGBoost model enables individualized AF risk assessment in middle-aged and older CHD patients, offering a practical tool for early identification and targeted intervention in clinical practice.
Feng Chen, Qin Fu, Ling Li et al.· Frontiers in Cardiovascular...· 0 citations
The CHA2DS2-VALa score significantly improves stroke risk stratification in AF by integrating LA diameter into conventional scoring, as echocardiographic measurement is widely available and reproducible.
S. Ömür, Emin Koyun, G. Genc Tapar et al.· Cardiovascular Electrophysio...· 0 citations