A Hybrid Ensemble Framework for Reliable and Interpretable Parkinson’s Disease Prediction Using Clinical Tabular Data
Abstract
Parkinson’s disease (PD) prediction in clinical tabular data is challenging due to feature interactions and data imbalance. The current methods put emphasis on either predictive performance or interpretability in a unified way. This paper presents a comparative hybrid learning framework that evaluates conventional machine learning models, boosting-based learners, deep learning (DL) architectures, and DL-augmented ensemble configurations for PD classification from clinical tabular data. The framework consists of feature selection, multiple base models, stacking-based hybrid ensembles, and explainable artificial intelligence (XAI) methods, including SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME), to improve clinical interpretability. Additional reliability is assessed through calibration analysis, robustness testing, bootstrap confidence intervals, subgroup analysis, and statistical validation. The experimental findings demonstrated strong internal performance across multiple evaluated model configurations. The experimental findings showed that model performance varied across model families and evaluation criteria. Among standalone boosting models, CatBoost achieved the best classification balance, with an accuracy of 0.9415, F1-score of 0.9527, and area under the curve (AUC) of 0.9710. Among hybrid ensemble configurations, combination 4 achieved the highest AUC of 0.9711, while combination 5 achieved the highest hybrid accuracy of 0.9399 and F1-score of 0.9517. The voting-based hybrid configuration also performed strongly, achieving an AUC of 0.9699. Among baseline machine learning models, Random Forest (RF) + Sequential Backward Elimination (SBE) achieved the highest baseline performance with an AUC of 0.9578, while TabTransformer was the robust DL model, achieving an accuracy of 0.8655, F1-score of 0.8895, and AUC of 0.9312. These results indicate stable internal performance within the evaluated dataset. However, because the experiments were conducted using a single publicly available dataset, external validation on independent clinical cohorts is required before real-world clinical deployment. To address the diagnostic circularity risk associated with the Unified Parkinson’s Disease Rating Scale (UPDRS), experiments were reported in two tracks: a full clinical assessment track including UPDRS and an UPDRS-excluded track for screening-oriented sensitivity analysis.