Aug 2026· RSC Advances· Vol 16, pp. 44973 - 44997· 0 citations· 43 references
Medicine
Abstract
Biodiesel production over metal-doped biochar and activated carbon (AC) catalysts involves complex nonlinear interactions among feedstock characteristics, catalyst descriptors, and operating conditions, making accurate yield prediction a challenging task. While machine learning (ML) has shown potential in process modeling, existing studies lack catalyst-aware frameworks that integrate material and process descriptors within a unified representation for biodiesel yield prediction. Furthermore, current approaches are often limited by small datasets and insufficient model interpretability, restricting their ability to support reliable catalyst screening and process optimization. To address these challenges, this study develops an explainable ML framework for biodiesel yield prediction using a literature-derived dataset of metal-doped biochar and AC catalyst systems. The framework integrates catalyst, feedstock, and operating-condition descriptors, augments sparse experimental data through curve digitization, evaluates six ML models, and applies SHAP and CatBoost-based explainability analysis. The neural network model achieved the highest predictive accuracy on unseen data, with RMSE of 3.27%, MAE of 1.64%, and R2 of 0.95, whereas linear regression showed the weakest performance, highlighting the nonlinear behavior of the catalytic system. Validation using an independent experimental dataset further confirmed model generalization. Explainability analysis identified alcohol-to-oil ratio, reaction time, catalyst amount, reaction temperature, and feedstock acid value as the key factors governing biodiesel yield. The proposed ML framework provides an accurate and interpretable approach for catalyst screening and data-driven optimization of sustainable biodiesel production processes.
This framework provides a tool for catalyst screening and process parameter optimization in biomass clean energy conversion systems and confirms that the model predicts syngas composition with less than five percentage points of absolute deviation.
Yadong Ge, Hongru Li, Zaixin Li et al.· Bioresource Technology· 0 citations
An interpretable Categorical Boosting model using native categorical encoding (CatBoost-NCE) was applied to a dataset containing 4026 experimental entries and provides a reliable tool for the data-driven discovery of heterogeneous catalysts.
Selective conversion of nitrogen-containing species into harmless molecular nitrogen (N2) remains a key challenge in the catalytic oxidation of nitrogen-containing volatile organic compounds (NVOCs). Machine learning (ML) provides an effective approach for predicting catalytic performance and identifying key descriptor...
Haotian Hu, Ying Wang, Zihao Zhai et al.· Journal of Colloid and Inter...· 1 citation
Accurately predicting bio-oil yield from biomass pyrolysis is a real challenge due to nonlinear interactions between feedstock physicochemical properties and operating conditions. In addition, high feature dimensionality, uncertainties in experimental measurements, and multicollinearity make prediction accuracy and int...
S. Almansour, L. Alkwai, Kusum Yadav et al.· Scientific Reports· 0 citations
Biomass-plastic catalytic co-pyrolysis offers a promising route to aromatics, yet high-value monocyclic aromatic (MAH) selectivity remains challenged by the complex interdependencies of reaction parameters. Herein, we present a machine learning-guided strategy to maximize MAH production from the co-pyrolysis of rice...
Meng-Ge Wu, Jie Li, Bo Jiang et al.· ACS Sustainable Chemistry &a...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.