Explainable Ensemble Learning Model Integrating Multidimensional Descriptors for CYP450 Inhibition Prediction of Kelulut Honey Phytochemicals
Abstract
Cytochrome P450 (CYP450) enzymes are involved in drug metabolism and pharmacokinetic behaviour and therefore prediction of CYP inhibition early in the drug discovery process is critical to reduce adverse drug reactions and late-stage drug failure. Earlier machine learning investigations achieved prediction accuracy of around 90% using ensemble learning algorithms such as Random Forest (RF) and Extreme Gradient Boosting (XGBoost). However, the majority of studies mostly relied on two dimensional (2D) molecular descriptors. In this study, an explainable ensemble machine learning workflow was developed using the KNIME Analytics Platform to predict CYP450 inhibition across five CYP isoforms using integrated RDKit-derived 2D descriptors and Weighted Holistic Invariant Molecular (WHIM) 3D descriptors. The results showed that the incorporation of 3D descriptors resulted in modest and model dependent changes compared to 2D descriptors alone, as there are structural differences among isoforms, and the 2D descriptors still dominated in feature importance. RF and XGBoost showed comparative performance, with test set AUC values ranging from 0.878 to 0.933 and external validation AUC values ranging from 0.741 to 0.963. SHAP analysis identified lipophilicity, molecular topology, ring systems, and molecular geometry as key contributors that enable a compound to interact and inhibit the CYP450 enzyme. Prediction of CYP450 enzyme on Kelulut honey phytochemicals showed reasonable agreement with SwissADME predictions. Overall, the integration of explainable machine learning with multidimensional molecular descriptors provides an effective and interpretable approach for early CYP450 liability assessment and compound prioritization in drug discovery.