Machine learning for prediction of newly diagnosed atrial fibrillation after emergency percutaneous coronary intervention during hospitalization in patients with acute ST-segment elevation myocardial infarction: a multi-center prospective study
Abstract
Background In-hospital complications after emergency PCI in acute STEMI are important, especially newly diagnosed atrial fibrillation and are associated with hemodynamic instability, heart failure, stroke and mortality. There are complex nonlinear interactions between clinical, inflammatory, cardiac remodeling and procedural factors that traditional risk scores may not adequately capture. Thus, this study aimed to develop and validate an interpretable machine-learning model for early prediction of in-hospital NDAF in patients with acute STEMI undergoing emergency PCI. Methods From January 2021 to December 2024, this study collected data on patients with acute STEMI after emergency PCI from the five tertiary general hospitals in Chongqing. Important clinical variables were identified using the selection operator and least absolute shrinkage. Based on the area under the curve, the best predictive model was selected from eight machine learning methods. The predictive model’s results were interpreted using Shapley Additive explanations. Results Eighteen variables were chosen for model building, and a total of 154 patients with acute STEMI following emergent PCI were included. With the largest area under the curve of 0.991 (95% CI: 0.963–1), sensitivity of 0.846, specificity of 0.914, and Brier score of 0.2356, the Gradient Boosting model was selected. The Shapley value analytical framework was systematically applied to decode feature importance patterns and illuminate individual prognostic determinants Conclusion The present study developed and compared eight machine learning models for the prediction of in-hospital NDAF among acute patients with STEMI treated with emergency PCI. The Gradient Boosting classifier achieved optimal predictive performance, highlighting its potential clinical utility for early risk stratification. By identifying high-risk individuals in advance, this model may support optimized clinical management and contribute to better prognosis in this high-risk population. Machine learning computations were implemented via a publicly accessible web server integrating the pre-trained Gradient Boosting prediction model. All analytical steps were performed following the platform’s default settings to maintain methodological standardization and reproducibility.