Machine learning-based modeling and analysis of solubility behavior in supercritical CO₂ systems under different operating conditions
Abstract
This study explores the analysis of Nystatin solubility and the density of supercritical carbon dioxide (SC-CO 2 ) in supercritical processing. A total of 28 experimental observations were initially collected, which were randomly divided into training (80%) and test (20%) subsets for model development and evaluation, respectively. Output variables include SC-CO 2 density and solubility of Nystatin, while input parameters include temperature and pressure. Four tree-based machine learning models, namely Random Forest (RF), Extremely Randomized Trees (ET), Gradient Boosting (GB), and XGBoost (XGB) were employed to predict these output variables. For hyper-parameter tuning, the Tabu Search (TS) algorithm was used. For the prediction of Nystatin solubility, the models exhibited commendable performance. Gradient Boosting (GB) outperformed others with an R 2 value of 0.98142, demonstrating a high level of accuracy in predicting solubility. It also achieved the lowest MAPE and RMSE, indicating superior predictive capabilities. In the case of SC-CO 2 density prediction, Random Forest (RF) and Extremely Randomized Trees (ET) models demonstrated strong performance with R 2 of 0.9375 and 0.95155, respectively. Overall, this research provides valuable insights into the estimation of solubility of Nystatin and the density of SC-CO 2 under varying temperature and pressure conditions. Specifically, machine learning models, specifically GB for solubility and RF and ET for density, has been demonstrated as valuable tools in the prediction of these significant properties.