Skip to content
Open access

Symbolic and domain-generalized machine learning for interpretable solubility modeling in supercritical CO₂.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

A domain-aware symbolic regression framework that discovers closed-form analytical expressions for ln(y), enabling interpretable and transferable solubility modeling across chemically diverse pharmaceutical compounds and demonstrates that domain-aware learning and symbolic discovery can jointly address the dual challenge of predictive robustness under domain shift and model interpretability in supercritical pharmaceutical solubility modeling.

Abstract

Accurate prediction of drug solubility in supercritical CO₂ remains challenging due to the limited generalizability of compound-specific correlations and the black-box nature of most machine learning models. This study proposes a domain-aware symbolic regression framework that discovers closed-form analytical expressions for ln(y), enabling interpretable and transferable solubility modeling across chemically diverse pharmaceutical compounds. A leave-one-drug-out (LODO) validation strategy is employed to rigorously assess extrapolation performance on unseen drug domains across a curated dataset of 196 experimental data points spanning 9 antihypertensive compounds. Under this evaluation, the domain-adversarial neural network (DANN) component of the proposed framework achieved RMSE = 0.33 ± 0.08, MAE = 0.24 ± 0.06, and R² = 0.89 ± 0.05, outperforming standard machine learning baselines including Random Forest, XGBoost, and Multi-Layer Perceptron in cross-domain generalization. The symbolic regression component additionally recovered compact closed-form analytical expressions that accurately represent the full experimental dataset (in-sample expression fit: R² = 0.962, RMSE = 0.031, MAE = 0.024) while remaining physically interpretable and analytically tractable. The combined framework demonstrates that domain-aware learning and symbolic discovery can jointly address the dual challenge of predictive robustness under domain shift and model interpretability in supercritical pharmaceutical solubility modeling.

Read PDF

Similar papers

Open access Aug 2026

A Leakage-Aware Benchmark Study of Machine Learning Models for Deep Eutectic Solvent Property Prediction

This study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.

Hakim Faraji, Julio Brito Santana, R. Rodríguez-Ramos et al. · 0 citations
#artificial intelligence Open access Sep 2026

Thermodynamic–molecular interdependencies and causal determinants governing pharmaceutical solubility in SC-CO2 systems

Accurate prediction of pharmaceutical solubility in supercritical CO₂ (SC-CO₂) systems is critical for green drug formulation and process intensification, yet existing machine learning studies largely prioritize predictive accuracy while overlooking mechanistic interpretability and causal understanding. This study pr...

Wael A. Mahdi, Adel Alhowyan, A. Obaidullah · 0 citations
Open access Sep 2026

Machine Learning for Toxicity Prediction in Low-Sample Molecular Classes

Deep learning models such as Chemprop have advanced quantitative molecular property prediction, but their reliance on large training sets limits use in data-scarce domains. We propose a framework that fine-tunes a general baseline model trained on publicly available data on small, class-specific datasets. The resulting...

Carlos Barajas, Laura L. Dunphy, Luke C. Mullany et al. · 0 citations
Open access Sep 2026

Comparative analysis using machine learning methods for determination of pharmaceutical solubility in supercritical CO2 considering physical properties of drugs

Background Reliable prediction of pharmaceutical solubility in supercritical carbon dioxide (SC-CO2) can support particle engineering and formulation design. However, record-wise validation may overestimate generalization when measurements for the same compound occur in both the training and test sets. Methods A leakag...

S. Alshahrani · 0 citations
Open access Aug 2026

Interpretable Small-Sample Deep Learning with Approximate Attribution Enhancement for Atomically Precise Gold Nanoclusters Synthesis

Small experimental datasets make synthesis-condition modelling particularly sensitive to overfitting and data leakage. Here, a graph convolutional neural network (GCNN) was coupled with approximate attribution enhancement (AAE) and evaluated on 54 gold-nanocluster synthesis records using a strict fold-local pipeline. A...

Zi-Yue You, Yi-Tong Qin, Yu-Bing Gao et al. · 0 citations
Open access Sep 2026

How effective are contrastive learning-based approaches for activity-cliff prediction?

Activity cliffs, defined as structurally similar molecules with vastly different properties represent a fundamental challenge in modern day drug discovery for property prediction models. While Graph Neural Networks (GNNs) have advanced molecular property prediction, they inherently struggle with this problem due to rep...

Akash Surendran, R. Miranda-Quintana · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.