Skip to content
Open access

Machine Learning- Based Cardiovascular Disease Risk Prediction in Hypertensive Patients: Explainable insights into Clinical Risk Factors

Aug 2026 · Nature Journal of Emerging Sciences Technologies and Innovations · 0 citations

TL;DR

Findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.

Abstract

Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.

Read PDF

Similar papers

Open access Aug 2026

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical deci...

T. Adeyemo · 0 citations
Sep 2026

Improving 10-year cardiovascular disease risk prediction using automated machine learning.

AIMS To develop a cardiovascular disease (CVD) risk prediction model with improved accuracy and interpretability by integrating diverse risk factors and applying Automated Machine Learning (AutoML), thereby enhancing clinical utility over conventional models. METHODS This is a prospective cohort study. Data were obta...

Si-Min He, Ju-Ping Wang, Le Zhao et al. · 0 citations
Open access Jul 2026

Explainable Ensemble Learning for Cardiovascular Risk Stratification A Multi-Hospital Stacking Approach with SHAP-Based Clinical Decision Support

This paper presents an explainable stacking ensemble framework for binary heart disease classification and three-tier risk stratification using multi-hospital cardiac data using a novel Dual-Balanced SMOTE strategy that simultaneously addresses class imbalance and gender bias.

B. Kusuma, Madhu M. Nayak · 0 citations
Open access Aug 2026

Machine Learning-Based Early Cardiovascular Disease Prediction: A Comparative Analysis of Supervised Learning Algorithms Using a Pakistani Clinical Dataset

The results confirm the effectiveness of ensemble learning approaches for medical diagnostics and highlight the potential of implementing ML-based screening tools in the limited resource healthcare environment in Pakistan.

Awais Khursheed, Soban Ahmed, Sibghat Ullah et al. · 0 citations
Open access Jul 2026

An interpretable machine learning model for predicting 1-year major adverse cardiovascular events in patients with type 2 diabetes and hypertension

An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN.

Juan Lv, Xi-Rui Wang, Zheng-Yi Zhang · 0 citations
Open access Jul 2026

Cross-Dataset Evaluation of Boosting Models for Hypertension Prediction

Hypertension remains a significant global risk factor for cardiovascular disease and related mortality, necessitating reliable early risk prediction models. Although boosting algorithms have demonstrated strong performance in structured medical data, limited studies have examined their consistency across heterogeneous...

Bety Wulan Sari, D. Murtiningsih, Donni Prabowo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.