Jul 2026· 2026 IEEE 27th China Conference on System Simulation Technology and its Applications (CCSSTA)· pp. 163-168· 0 citations· 18 references
Abstract
Drug-induced liver injury (DILI) is a major cause of drug development failure and post-marketing withdrawal. Accurate computational prediction of hepatotoxicity is hindered by complex biological mechanisms and scarce labeled toxicity data. Although pretrained molecular language models like ChemBERTa perform well in molecular property prediction, their generalization ability for small-sample DILI prediction remains underexplored. Here, we systematically compared traditional molecular fingerprint-based machine learning methods and ChemBERTa-based models for DILI classification on the DILIst dataset. Canonical SMILES from PubChem were used to generate Morgan fingerprints and ChemBERTa embeddings. We evaluated Random Forest, XGBoost, full fine-tuning, frozen encoder transfer learning, and embedding-based classifiers under both random and scaffold data splits. Results showed that Morgan fingerprints combined with Random Forest achieved the best performance, with ROC-AUC of 0.783 and PR-AUC of 0.849 under random split. Scaffold split markedly degraded the performance of all models, indicating poor generalization to unseen chemical scaffolds. ChemBERTa embedding-based classifiers outperformed end-to-end fine-tuning, suggesting that pretrained representations are better used as fixed feature extractors under limited labeled DILI data. Further SHAP analysis detected key toxicity-related molecular fragments, and t-SNE showed insufficient latent-space separation between DILI-positive and negative compounds. Our results confirm that traditional fingerprint-based machine learning remains highly competitive for small-sample hepatotoxicity prediction, and this work provides a reliable computational framework for early drug safety assessment.
Drug-induced ocular toxicity is difficult to predict and evaluate, particularly for prostaglandin F2α (PGF2α) analogs used in the management of glaucoma. Traditional global predictive models, which are built with large and various datasets (n = 6187), provide systematic false-negative results for groups of structurally...
Xin-Yi Lu, Li Ren, Chen Wang et al.· Molecules· 0 citations
Accurate prediction of physicochemical properties such as the octanol-water partition coefficient’s logarithm (LogP) is critical in early-stage drug development. This study presents a novel and interpretable computational framework that integrates symbolic graph-theoretical descriptors, specifically topological indices...
Shabbir Ahmad, Sana Javed, S. Khalid et al.· Journal of King Saud Univers...· 0 citations
This study systematically benchmark six ranking loss functions, including state-of-the-art listwise methods, and five types of molecular representations across two large-scale drug screening datasets, CTRP and PRISM, to demonstrate that listwise loss functions such as LambdaLoss and LambdaRank consistently excel in bot...
Faraz Sarmeili, Benyamin Ghahremani-Nezhad, Mohammad Khalilpour et al.· PLoS ONE· 0 citations
Traditional QSAR toxicity models are, in general, assessed with random train-test splits where structural overlaps between train and test compounds are allowed. This usually results in an inflated predictive performance. Therefore, this paper looks at the problem of toxicity prediction under the structural distribution...
Wael A. Mahdi, Adel Alhowyan, A. Obaidullah· Scientific Reports· 0 citations
Binding affinity estimation is an important computational task in modern drug discovery because it helps prioritize compounds according to their expected interaction strength with protein targets. This study presents CAMELEON-AI X v3.0, a machine-learning framework designed to improve binding affinity prediction on het...
Sahil Dipak Lonkar· International Journal For Mu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.