EXPLAINABLE MACHINE LEARNING FOR PREDICTING MOLECULAR TOXICITY FROM PHYSICOCHEMICAL AND STRUCTURAL PROPERTIES
Abstract
The prediction of molecular toxicity is becoming more and more critical for effective chemical safety assessment, drug development and environmental risk assessment. In this study, an explainable machine learning model for predicting molecular toxicity was developed using physicochemical descriptors and structural properties of Tox21 dataset. There were 7,831 molecules and 12 binary toxicity endpoints in the dataset, with 7,823 molecules remaining after structure validation for the molecules. Morgan molecular fingerprints were created along with physicochemical descriptors such as molecular weight, LogP, topological polar surface area, hydrogen-bond properties, rotatable bonds, heavy-atom count, ring properties, and Fraction Csp3. The accuracy, precision, recall, specificity, F1 score, balanced accuracy, ROC-AUC and PR-AUC were used to evaluate the following logistic regression, random forest, XGBoost and Support Vector Machine models. Random Forest had the best mean ROC-AUC (0.839) and PR-AUC (0.480) values while XGBoost had the best balanced accuracy (0.738) and F1 score (0.402). Generally, active molecules were larger, more lipophilic and more aromatic than inactive compounds for the representative SR-ARE endpoint. The molecular weight, LogP, Fraction Csp3, TPSA, number of heavy atoms and aromatic properties were found to be influential predictors using SHAP analysis. In general, this work shows that explainable machine learning offers accurate and interpretable molecular toxicity predictions to enable compound prioritization and computational chemical-safety assessment.