Skip to content
Open access

A Hybrid GA-KNN Framework For Cardiovascular Disease Prediction Using Optimized Clinical Feature Selection

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 38 references

TL;DR

An optimized hybrid approach of the genetic algorithm and K-nearest neighbor method for cardiovascular disease prediction is proposed and it is demonstrated that optimized GA-KNN can achieve both feature dimensions for the initial screening of cardiovascular diseases.

Abstract

Background Study: Background Study: Cardiovascular disease (CVD) causes millions of fatalities each year and places a heavy financial strain on healthcare systems. Better patient outcomes, prompt clinical intervention, and lower healthcare costs all depend on early and precise cardiovascular disease prediction. Through the analysis of massive amounts of clinical data, machine learning algorithms have considerable potential in helping doctors identify diseases. Problem Statement: High-dimensional clinical datasets, repetitive and irrelevant features, and the difficulty to consistently identify the most discriminative risk factors are common problems for current machine learning-based techniques for cardiovascular disease prediction. These problems limit the robustness and generalizability of prediction models, raise computing costs, and decrease classification accuracy. This is particularly true for distance-based classifiers, such as K-Nearest Neighbor (KNN). Developing an efficient approach that combines accurate classification with suitable clinical feature selection remains a critical research problem for improving early cardiovascular disease prediction and enabling reliable clinical decision-making. Purpose: In a medical decision support system, the prediction of cardiovascular disease is an important task, as early detection can help in minimizing the risk of mortality, delay in treatment, and cost of healthcare. Methods: In this study, an optimized hybrid approach of the genetic algorithm and K-nearest neighbor method for cardiovascular disease prediction is proposed. The clinical attributes are selected using the genetic algorithm, and the final classifier is KNN. Four datasets, the Cleveland Processed Heart Dataset, the CRPF Ranchi Clinical Heart Dataset, the Cleveland Hungarian Statlog Dataset, and the Heart Failure Clinical Record Dataset, were used for evaluating the model. Initial experiments were conducted with k-fold values of 5, 10, 15, 20, and 25 folds, and then an optimized 10-fold GA-KNN approach with feature selection, normalization, binary target conversion, and hyperparameter tuning of KNN was executed. Results: The optimized model achieved accuracies of 78.19%, 75.71%, 92.10%, and 81.98%, respectively, with ROC-AUC values of 0.8612, 0.7603, 0.9665, and 0.8313. Conclusion: It is demonstrated that optimized GA-KNN can achieve both  feature dimensions for the initial screening of cardiovascular diseases. The proposed GA-KNN framework is simple, interpretable, and computationally efficient for preliminary cardiovascular disease screening.

Read PDF

Similar papers

Review Open access Aug 2026

Artificial Intelligence for Early Heart Disease Prediction: A Review of Machine Learning Techniques

There is an urgent need for explainable, clinically validated and standardised ML frameworks to translate predictive models into routine healthcare practice and improve early detection of cardiovascular disease.

Hanna Rasheed, Arya.K.R Arya.K.R, Ashida.K.A Ashida.K.A · 0 citations
#explainable ai Review Open access Sep 2026

Comparative Analysis of Artificial Intelligence Techniques for Cardiovascular Disease Diagnosis and Risk Prediction

Evidence is provided that ensemble-based frameworks currently offer the most effective balance between predictive accuracy, robustness, and clinical feasibility, and future research should emphasize multi-center external validation and explainable AI frameworks.

Marium Shaikh, Hanmant Fadewar · 0 citations
Open access Aug 2026

Enhancing Heart Disease Prediction Through The Hybrid Random Forest–Gradient Boosting-Logistic Regression Model (HRFGLM): A Data-Driven Predictive Framework

This study suggests a methodology for early cardiovascular disease prediction using various machine learning techniques for various prediction objectives, and implemented a few models, which include Gradient Boost, Random Forests, and Linear Regression classifiers getting 75.19% accuracy.

Asha Dilipkumar Jariwala, Hemangini G. Patel · 0 citations
Open access Aug 2026

Explainable AI-Driven Decision Support System for Early Prediction of Cardiovascular Diseases

An Explainable AI-Driven Decision Support System for the early prediction of cardiovascular diseases by integrating intelligent clinical data preprocessing, feature selection, an ensemble learning-based prediction model, and explainable artificial intelligence is proposed.

K. Sridhar, S. Swathi, S. Saranya et al. · 0 citations
Conference Aug 2026

Gradient Boosting Techniques in a Risk-Aware and Explainable Machine Learning Framework for Heart Disease Prediction

The complicated connection between medical risk factors and the serious ramification of misdiagnosis highlights the essential challenge of detecting cardiovascular disease in its early stages. Although most examinations to date have focused on accuracy-centric evaluation, which may not absolutely account for clinical s...

H. Suresh, P. R. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.