Skip to content
Open access

An Explainable Ensemble Machine Learning Framework for PCOS Risk Prediction Using Clinical and Hormonal Data

Jul 2026 · Turkish Journal of Engineering · Vol 10, pp. 864-874 · 0 citations · 23 references

TL;DR

The proposed explainable ensemble framework presents a scalable, accurate, and interpretable decision support system for PCOS that is feasible to adopt in practical healthcare environments, especially in rural areas where medical resources are limited and more medical aids for the detection of such diseases are desperately needed.

Abstract

Polycystic Ovary Syndrome (PCOS) is the most prevalent endocrine disorder affecting women of reproductive age, and there is a need for early and accurate diagnosis to prevent long-term reproductive and metabolic consequences. This work introduces a transparent ensemble model to predict PCOS at the individual level using everyday clinical, hormonal, metabolic, and lifestyle parameters. This work allows patient-based prediction, as this includes biologically plausible predictors such as serum testosterone, the luteinizing hormone to follicle-stimulating hormone ratio (LH/FSH), insulin, and menstrual irregularities. A scheme for data preprocessing, feature selection, model training, and testing is proposed. The performance of four classical classifiers, i.e., Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Gradient Boosting (GB), is evaluated. Then, a voting-based ensemble method is proposed to enhance robustness and generalization. Results of experiments confirm good predictive performance, even for a significantly challenging task, with 96% accuracy and an ROC–AUC of 0.99, while decreasing false-negative rates is highly important for early screening. To add a layer of transparency and clinical trustworthiness, SHAP-based explainable artificial intelligence was adopted to evaluate global and patient-level feature importance. In addition to binary prediction, the proposed model reduces risk stratification to a probabilistic scale (low, moderate, high), making it more practical for clinical decision support. In conclusion, our proposed explainable ensemble framework presents a scalable, accurate, and interpretable decision support system for PCOS that is feasible to adopt in practical healthcare environments, especially in rural areas where medical resources are limited and more medical aids for the detection of such diseases are desperately needed. It has strong potential for integration into AI-enabled clinical screening systems.

Read PDF

Similar papers

#explainable ai Open access Aug 2026

A Two‐Stage Hybrid Feature Selection and Ensemble Learning Framework With Explainable AI for Accurate PCOS Prediction

Combining hybrid ensemble learning, two-stage feature selection, and XAI approaches provides a computationally efficient, dependable, and interpretable method for PCOS diagnosis and practitioners may find this model to be a useful decision-support tool that improves the accuracy of diagnosis and lessens the need for human interpretation.

Md Rakibul Hasan Efty, M. Rohman, K. M. Uddin et al. · 0 citations
Open access Aug 2026

Automated PCOS Disease Detection Using Clinical and Diagnostic Features

Findings indicate that explainable machine learning models, particularly KNN and XGBoost, provide accurate and interpretable decision support for early PCOS screening, enabling timely intervention and offering a promising foundation for intelligent healthcare decision-support systems.

Sana Rubab, Musarrat Shaheen, Zohrain Tabassum et al. · 0 citations
Conference Aug 2026

PCOS-XAI: An Interpretable Predictive Model for Polycystic Ovary Syndrome Using Feature Selection and Ensemble Classification

Polycystic Ovary Syndrome (PCOS) is a disease that has spread across the globe and has become a significant health concern that mainly affects women of reproductive age. The detection, diagnosis, treatment, and management of the condition at an early stage are vital in order to lower the risk of long-term complications, primarily an elevated risk of type 2 diabetes and gestational diabetes. In line with advancements in computational methods, machine learning and ensemble learning techniques have drawn significant attention as a means of facilitating automated medical diagnosis. This paper aims to build a reliable, efficient, and interpretable PCOS diagnostic system that not only supports evidence-based decision-making but also provides global explanations of feature contributions. To identify the best-performing model and reduce the number of features, six machine learning algorithms-Logistic Regression, Random Forest, Decision Tree, Naive Bayes, Support Vector Machine, and k-Nearest Neighbors-have been employed, with Bayesian hyperparameter tuning applied for optimization. High-performing base learners, together with a meta-learner, were then stacked using an ensemble approach to further enhance predictive performance. Experiments were conducted on a publicly available PCOS dataset using two train-test split ratios (70:30 and 80:20). The results show that the stacking ensemble model along with optimized features demonstrated improved performance over individual baseline models.

Pooja Snehal Janwe, Nazia Nusrath Ul Ain, K. Radhika et al. · 0 citations
Aug 2026

AI meets Reproductive Health: Early Diagnosis of PCOS using Machine Learning

Polycystic Ovary Syndrome (PCOS) is a common complex hormonal condition that disrupts the balance in metabolism, fertility, and dermatological health, especially among women of reproductive age. It's mixed, and superimposing clinical analysis often tends to slow down the correct diagnosis. However, in the recent past, the convergence of data science and healthcare has made it possible to analyse large-scale patient data to gain a deeper understanding of complex diseases such as PCOS. The suggested methodology implies the use of a multi-dimensional database comprising physiological measurements, hormone concentration, lifestyle factors, and clinical indicators to find trends that can be related to the condition. A data-driven systematic method is used to examine statistical relations and hidden patterns. Pre-processing of data (dealing with missing values, outlier detection, normalising values and categorical encoding) is done, followed by the Exploratory Data Analysis (EDA), which reveals the major associations and visual representations of the distributions of variables. The feature engineering adds new features like hormone ratios, BMI groups and waist to hip measurements to enhance model performance. Improved oversampling methods, namely ADASYN and Borderline-SMOTE, reduce the imbalance between classes, and state-of-the-art models of machine learning, such as Logistic Regression, Random Forest, XGBoost, CatBoost, and ensemble voting classifiers, are applied to the PCOS classification. Clinical relevance is statistically tested (Chi-square, t-tests, ANOVA), and the explainable AI using SHAP is used to complement model interpretability. The findings aid in describing the significance of the machine learning methodology to be supplemented by the analytical methods to produce useful clinical data, enhance the quality of diagnosis, and promote the initial PCOS detection, thus contributing to the general achievement of AI in the field of prophylaxis and personalised medicine.

Abhinav Pathak, M. Sujithra, H. P. et al. · 0 citations
Open access Aug 2026

Comparative machine learning analysis identifies random forest and adaboost as superior models for the evaluation of semen quality and reproductive hormones

Background Semen analysis is a widely accepted laboratory investigation for evaluating male infertility. However, in cases of idiopathic or unexplained male infertility, comprehensive assessment may require blood-based profiling of a reproductive hormone panel—including follicle-stimulating hormone (FSH), luteinizing hormone (LH), prolactin (PRL), and testosterone—to clearly identify endocrine abnormalities that may contribute to infertility. These variables exhibit nonlinear dynamics and multidimensional interdependencies, presenting analytical challenges for conventional statistical approaches to detect subtle, latent patterns within such complex data. In contrast, advanced analytics such as machine learning (ML) algorithms enable robust modelling and interpretation of high-dimensional datasets. Methodology Pre-processing of data included normalization and correlation analysis. Using a train-test split ratio of 80:20, nine ML algorithms—Linear Regression, Lasso Regression, Ridge Regression, Elastic Net, Random Forest, Support Vector Regression, Gradient Boosting, AdaBoost, and Neural Networks—were trained using cross-validation and hyperparameter optimization. Model performance was assessed using the coefficient of determination (R2), root mean square error (RMSE), and feature importance analysis. Results Ensemble learning approaches consistently outperformed conventional regression models. AdaBoost achieved the highest predictive accuracy for sperm motility (R2 = 0.993), while Gradient Boosting yielded superior predictions for progressive motility (R2 = 0.932) and vitality (R2 = 0.982). Random Forest demonstrated the strongest performance for semen volume (R2 = 0.428), LH (R2 = 0.432), and PRL (R2 = 0.586). Conversely, predictions for seminal pH (R2 = 0.037), liquefaction time (R2 = 0.158), FSH (R2 = 0.178), and testosterone (R2 = −0.114) were limited. Feature importance analysis identified total sperm count, sperm concentration, non-progressive motility, morphology, and non-motile sperm percentage as the most influential predictors across models. Conclusions Random Forest and AdaBoost emerged as the most effective and broadly applicable models for evaluating male reproductive parameters, whereas Gradient Boosting exhibited exceptional predictive capacity for select semen parameters.

S. Roychoudhury, S. Paul, Birupakshya Paul Choudhury et al. · 0 citations
Open access Aug 2026

Methodology for a Dual-Target Reproductive Intelligence System for Early Prediction of Infertility and Menopause

Reproductive health of women involves complex, heterogeneous and interdependent clinical factors that make early risk assessment difficult through manual valuation alone. This study proposes a dual-target machine-learning reproductive intelligence system for predicting infertility risk and menopause transition from shared health data. The system integrates clinical records, hormonal profiles, laboratory results, ultrasound findings, menstrual history, age, body mass index, lifestyle factors, infection history and patient-reported symptoms. It addresses a limitation in existing reproductive-health applications which commonly focus on isolated tasks, such as ovulation tracking, embryo grading, assisted-reproduction outcomes, symptom monitoring, without jointly modelling infertility and menopause across the reproductive life course. The proposed architecture uses a stacked ensemble of Extreme Gradient Boosting, Extra Trees and Convolutional Neural Networks as base learners. Their predictions are combined via a Random Forest meta-learner to classify women into infertility-risk categories and menopause-transition stages. This design exploits complementary strengths in nonlinear feature learning, variable interaction detection and robust classification. SHAP additive explanations are incorporated to identify the contribution of each predictor and provide patientlevel and global explanations. Thereby, improving transparency and clinical interpretability. The system is intended to support earlier screening, risk stratification, referral, and personalised reproductive-health decisions rather than replace professional diagnosis. Its application is especially relevant in Nigeria and similar African settings where reproductive data remain underused and delayed assessment, repeated hospital visits, stigma, emotional distress and limited menopause services persist. The proposed methodology provides an integrated, explainable and context-sensitive foundation for intelligent decision support across the reproductive ageing continuum in women

Folayemi Faith Adekola, Oyebode Aduragbemi, Olufunke Olubukola Ayennakin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.