Predicting the Classification of Heart Failure Patients Using Optimized Machine Learning Algorithms
Abstract
Heart failure prediction is a critical task in healthcare analytics, enabling early diagnosis and timely intervention to reduce mortality rates. However, traditional clinical approaches often lack scalability and struggle to capture complex nonlinear relationships in patient data. To address these limitations, a machine learning-driven framework is developed using the healthcare-dataset-stroke-data, comprising diverse clinical attributes associated with cardiovascular conditions. The data undergoes preprocessing steps including duplicate removal, missing value imputation, label encoding, normalization via StandardScaler, genetic algorithm-based feature selection, and class balancing using SMOTE. Multiple models are implemented, including Support Vector Machine, Random Forest, and Gradient Boosting with GridSearch, Bayesian Optimization, and Particle Swarm Optimization, along with ensemble methods such as Voting and Stacking classifiers. Model performance is evaluated using accuracy, precision, recall, and F1-score, supported by confusion matrix analysis. Results show that Stacking and Voting classifiers achieve the highest performance of 0.971 across all metrics, outperforming optimized individual models, while Gradient Boosting with PSO performs best among single models. Furthermore, interpretability is enhanced using LIME and SHAP techniques, and a Flask-based interface enables real-time prediction deployment. This integrated approach improves both prediction accuracy and model transparency in heart failure classification.