Skip to content
Open access

Machine Learning for Toxicity Prediction in Low-Sample Molecular Classes

Sep 2026 · bioRxiv · 0 citations · 25 references
Biology

Abstract

Deep learning models such as Chemprop have advanced quantitative molecular property prediction, but their reliance on large training sets limits use in data-scarce domains. We propose a framework that fine-tunes a general baseline model trained on publicly available data on small, class-specific datasets. The resulting models retain the baseline’s generalization ability while gaining class-specific accuracy and produce probabilistic outputs that capture uncertainty in the training data. We demonstrate the approach on three toxicity classes defined by a common core structure, target, or mode of action: (i) organophosphates, (ii) androgen receptor antagonists, and (iii) estrogen receptor β antagonists. Each fine-tuned model outperforms classical machine-learning methods and the EPA TEST tool. The probabilistic nature of the predictions enables prioritization of compounds for experimental validation and seamless integration with data streams of varying quality, supporting iterative decision-making in chemical safety and drug discovery.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.