A domain-aware symbolic regression framework that discovers closed-form analytical expressions for ln(y), enabling interpretable and transferable solubility modeling across chemically diverse pharmaceutical compounds and demonstrates that domain-aware learning and symbolic discovery can jointly address the dual challenge of predictive robustness under domain shift and model interpretability in supercritical pharmaceutical solubility modeling.
Abstract
Accurate prediction of drug solubility in supercritical CO₂ remains challenging due to the limited generalizability of compound-specific correlations and the black-box nature of most machine learning models. This study proposes a domain-aware symbolic regression framework that discovers closed-form analytical expressions for ln(y), enabling interpretable and transferable solubility modeling across chemically diverse pharmaceutical compounds. A leave-one-drug-out (LODO) validation strategy is employed to rigorously assess extrapolation performance on unseen drug domains across a curated dataset of 196 experimental data points spanning 9 antihypertensive compounds. Under this evaluation, the domain-adversarial neural network (DANN) component of the proposed framework achieved RMSE = 0.33 ± 0.08, MAE = 0.24 ± 0.06, and R² = 0.89 ± 0.05, outperforming standard machine learning baselines including Random Forest, XGBoost, and Multi-Layer Perceptron in cross-domain generalization. The symbolic regression component additionally recovered compact closed-form analytical expressions that accurately represent the full experimental dataset (in-sample expression fit: R² = 0.962, RMSE = 0.031, MAE = 0.024) while remaining physically interpretable and analytically tractable. The combined framework demonstrates that domain-aware learning and symbolic discovery can jointly address the dual challenge of predictive robustness under domain shift and model interpretability in supercritical pharmaceutical solubility modeling.
This study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.
Hakim Faraji, Julio Brito Santana, R. Rodríguez-Ramos et al.· ACS Omega· 0 citations
Accurate prediction of pharmaceutical solubility in supercritical CO₂ (SC-CO₂) systems is critical for green drug formulation and process intensification, yet existing machine learning studies largely prioritize predictive accuracy while overlooking mechanistic interpretability and causal understanding. This study pr...
Wael A. Mahdi, Adel Alhowyan, A. Obaidullah· Scientific Reports· 0 citations
Deep learning models such as Chemprop have advanced quantitative molecular property prediction, but their reliance on large training sets limits use in data-scarce domains. We propose a framework that fine-tunes a general baseline model trained on publicly available data on small, class-specific datasets. The resulting...
Carlos Barajas, Laura L. Dunphy, Luke C. Mullany et al.· bioRxiv· 0 citations
Background Reliable prediction of pharmaceutical solubility in supercritical carbon dioxide (SC-CO2) can support particle engineering and formulation design. However, record-wise validation may overestimate generalization when measurements for the same compound occur in both the training and test sets. Methods A leakag...
Small experimental datasets make synthesis-condition modelling particularly sensitive to overfitting and data leakage. Here, a graph convolutional neural network (GCNN) was coupled with approximate attribution enhancement (AAE) and evaluated on 54 gold-nanocluster synthesis records using a strict fold-local pipeline. A...
Activity cliffs, defined as structurally similar molecules with vastly different properties represent a fundamental challenge in modern day drug discovery for property prediction models. While Graph Neural Networks (GNNs) have advanced molecular property prediction, they inherently struggle with this problem due to rep...
Akash Surendran, R. Miranda-Quintana· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.