Aug 2026· International Journal of Innovative Science and Research Technology· 0 citations· 21 references
TL;DR
This study focuses on leveraging Random Forest model for PCOS prediction using only clinical and biochemical data, enhanced with Shapley Additive Explanations (SHAP) for model interpretability.
Abstract
Polycystic Ovary Syndrome (PCOS) is a lead major disorder and primary cause of infertility in women of
reproductive age, affecting about 13% of this population globally with over 70% of cases remaining undiagnosed. Early
diagnosis is yet challenging, particularly in low-resource settings where ultrasound imaging is inaccessible. This study
focuses on leveraging Random Forest (RF) model for PCOS prediction using only clinical and biochemical data, enhanced
with Shapley Additive Explanations (SHAP) for model interpretability. A publicly available Kaggle PCOS dataset from
541 women (177 PCOS-positive, 364 negative) across 10 hospitals in Kerala, India, was utilised. A leakage-free
preprocessing pipeline applied median and mode imputation before data splitting.
Polycystic Ovary Syndrome (PCOS) is a disease that has spread across the globe and has become a significant health concern that mainly affects women of reproductive age. The detection, diagnosis, treatment, and management of the condition at an early stage are vital in order to lower the risk of long-term complications, primarily an elevated risk of type 2 diabetes and gestational diabetes. In line with advancements in computational methods, machine learning and ensemble learning techniques have drawn significant attention as a means of facilitating automated medical diagnosis. This paper aims to build a reliable, efficient, and interpretable PCOS diagnostic system that not only supports evidence-based decision-making but also provides global explanations of feature contributions. To identify the best-performing model and reduce the number of features, six machine learning algorithms-Logistic Regression, Random Forest, Decision Tree, Naive Bayes, Support Vector Machine, and k-Nearest Neighbors-have been employed, with Bayesian hyperparameter tuning applied for optimization. High-performing base learners, together with a meta-learner, were then stacked using an ensemble approach to further enhance predictive performance. Experiments were conducted on a publicly available PCOS dataset using two train-test split ratios (70:30 and 80:20). The results show that the stacking ensemble model along with optimized features demonstrated improved performance over individual baseline models.
Pooja Snehal Janwe, Nazia Nusrath Ul Ain, K. Radhika et al.· International Conference on...· 0 citations
Combining hybrid ensemble learning, two-stage feature selection, and XAI approaches provides a computationally efficient, dependable, and interpretable method for PCOS diagnosis and practitioners may find this model to be a useful decision-support tool that improves the accuracy of diagnosis and lessens the need for human interpretation.
Md Rakibul Hasan Efty, M. Rohman, K. M. Uddin et al.· Health Science Reports· 0 citations
Polycystic Ovary Syndrome is a common endocrine disorder characterized by ovulatory dysfunction, hyperandrogenism, and/or polycystic ovarian morphology, with significant reproductive and metabolic consequences. Due to heterogeneous symptom profiles, Polycystic Ovary Syndrome is frequently underdiagnosed or diagnosed late. In this study, we develop machine learning models for early Polycystic Ovary Syndrome prediction using a structured clinical dataset with 42 features and 542 patient records. After data cleaning and normalization, correlation-based feature selection was applied to retain the most predictive variables. Multiple models were trained and evaluated, including Logistic Regression, Decision Tree, KNN, and Random Forest. Results demonstrate that Random Forest achieves the best overall performance (approximately 88% accuracy), suggesting that ensemble models can effectively capture non-linear feature interactions in clinical data. We also contextualize findings with international clinical guidance and recent work on explainable and clinically applicable Polycystic Ovary Syndrome prediction systems.
D. Thota, A. Mahesha, M. F. Khasim et al.· medRxiv· 0 citations
SHAP (SHapley Additive exPlanations) analysis supports the assertion that the quantity of follicles, FSH/LH ratio, and the level of LH should be considered the most prevalent predictors, and that the evidence provided by the analysis can be interpreted by clinicians and corresponds to the Rotterdam diagnostic criteria.
Arya Malode, Vijayshri A. Injamuri, Vikul J. Pawar et al.· International journal of com...· 0 citations
Polyendocrine metabolic ovarian syndrome (PMOS) previously known as Polycystic Ovary Syndrome (PCOS) is a common endocrine disorder affecting women of reproductive age, yet its diagnosis remains challenging because accurate assessment often depends on investigations that are costly, less accessible, and difficult to scale in resource-constrained clinical settings. We develop and assess a cost-aware modelling framework for PCOS risk stratification, comparing an inexpensive-feature model based on demographic, symptom, anthropometric, and routine clinical variables with an augmented model that additionally incorporates laboratory and ultrasound features when fuller diagnostic work-up is available. Logistic Regression (LR) and Random Forest (RF) models are evaluated using discrimination and calibration metrics, including AUC, accuracy, F1, precision, recall, Brier score, and expected calibration error (ECE), alongside clinical utility through decision curve analysis (DCA). To support safer decision-making under constrained capacity, we further integrate conformal prediction to provide finite-sample coverage guarantees and controlled abstention. Across train-test and out-of-fold evaluations, the augmented model showed consistent performance gains over the inexpensive-feature model, with AUC increasing by 6.9% for LR and 7.4% for RF. LR exhibited more favourable calibration in most settings, while RF achieved higher precision. Decision curve analysis showed higher net benefit for the augmented model across clinically relevant threshold regions, although feature sensitivity analysis indicated that a compact subset of inexpensive predictors preserved over 80% of maximal AUC, supporting the practical value of lower-burden screening. Conformal prediction achieved 94.5% overall coverage with a 41.3% abstention rate, maintaining near-nominal validity across age and BMI subgroups. Under constrained referral capacity, prioritising highest-risk cases maximised net benefit at lower capacity levels, while uncertainty-aware deferral became more compatible with decision utility as available capacity increased. These findings support a proof-of-concept framework in which inexpensive information provides meaningful early risk stratification, while added laboratory and ultrasound inputs offer incremental value when more resource-intensive assessment is feasible.
A. Adegoke, Idris Babalola, P. A. Odesola· Journal of Engineering Resea...· 0 citations
Findings indicate that explainable machine learning models, particularly KNN and XGBoost, provide accurate and interpretable decision support for early PCOS screening, enabling timely intervention and offering a promising foundation for intelligent healthcare decision-support systems.
Sana Rubab, Musarrat Shaheen, Zohrain Tabassum et al.· Biomedical Informatics and S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.