Skip to content
Open access

An Explainable Machine Learning Framework for Hypothesis Generation in Biochemical Methane Potential Prediction

Aug 2026 · Bioenergy Research · Vol 19 · 0 citations · 26 references

TL;DR

The framework may support preliminary feedstock screening, prioritization of confirmatory experiments, and hypothesis generation within the observed compositional domain, but it is not intended for causal inference or field-ready engineering prediction.

Abstract

Biochemical methane potential (BMP) assays are widely used to evaluate feedstocks for anaerobic digestion, but they are slow and resource-intensive. This study evaluated whether explainable machine learning can support transparent screening and hypothesis generation from compositional feedstock data. A Random Forest model was trained on 127 solid and semi-solid feedstocks from a public dataset (Dry Matter, DM ≥ 15%). For the primary BMP per dry matter setting, repeated nested five-fold cross-validation yielded R2 = 0.530, RMSE = 44.5 Nm3 CH4/t DM, and MAE = 33.6 Nm3 CH4/t DM; target-sensitivity analyses gave lower performance when DM and VS were excluded from the predictor set or when BMP per VS was modeled directly, with R2 values of 0.423 and 0.377, respectively. The evaluated models showed broadly similar, moderate performance with the available dataset and descriptor set. Lignin remained a leading predictor across Random Forest, Gradient Boosting, and Extra Trees, whereas the ordering of other leading descriptors was model-dependent. SHAP visualization and stratified analyses suggested a possible DM-stratified Lignin-BMP pattern, but the adjusted interaction was not significant and bootstrap uncertainty spanned zero; therefore, this pattern was treated as exploratory. Accordingly, the framework may support preliminary feedstock screening, prioritization of confirmatory experiments, and hypothesis generation within the observed compositional domain, but it is not intended for causal inference or field-ready engineering prediction.

Read PDF

Similar papers

Review Aug 2026

Predicting methane production in biochar-assisted anaerobic digestion using random forest and SHAP analysis of multi-study experimental data.

Anaerobic digestion (AD) optimization for enhanced biogas production remains challenging because methane production is governed by complex interactions among substrate characteristics, biochar (BC) properties, and operating conditions. In this study, a literature-derived dataset comprising 623 experimental conditions f...

L. Huaraca, C. Almeida-Naranjo, Paul Sarango-Lalangui · 0 citations
Open access Sep 2026

PS6-3. Enhancing Methane Emissions Modeling Through Feature Engineering and Machine Learning.

Enteric methane produced by beef cattle is a major greenhouse gas and represents a loss of feed energy, making accurate prediction of methane emissions an important step toward developing effective mitigation strategies and improved production performance. This study evaluated the effectiveness of predictive models f...

Serinmary Pulikkottil Rejimon, Sreekar Veeranki, Aushmeet Singh et al. · 0 citations
Open access Aug 2026

Valorizing Residue Biomass into Bioenergy: An Explainable Hybrid Machine Learning Model for Predicting Higher Heating Value (HHV) from Elemental Composition

Transforming waste and agricultural-residue biomass into bioenergy is central to the circular bioeconomy, yet routing such heterogeneous residues to the right thermochemical pathway depends on the higher heating value (HHV), which is conventionally measured by slow, resource-intensive bomb calorimetry. Here, we present...

Y. Özüpak, Emrah Aslan, Mehmet Burukanli et al. · 0 citations
Aug 2026

A machine learning-based dual system coupling model of catalytic-gasification: a calcium-based case study with mechanistic interpretability.

This framework provides a tool for catalyst screening and process parameter optimization in biomass clean energy conversion systems and confirms that the model predicts syngas composition with less than five percentage points of absolute deviation.

Yadong Ge, Hongru Li, Zaixin Li et al. · 0 citations
Open access Aug 2026

A robust framework integrating random standard deviation sampling with machine learning for volatile fatty acids prediction in Fe3O4-mediated high-load anaerobic fermentation.

Predicting volatile fatty acids (VFAs) in Fe3O4-mediated high-load anaerobic fermentation (AF) is hindered by stochastic uncertainty and intrinsic data scarcity. To bridge the gap between the lab-scale datasets and industrial prediction demands, this study established the RSDS-ML framework integrating random standard d...

Lanting Wang, Tengshuang Ma, Ling Deng et al. · 0 citations
Open access Sep 2026

Machine Learning-Based Prediction of N2O Emissions from Tea Plantations and Identification of Driving Factors for Sustainable Nitrogen Management

The GBRT-based model provides a useful tool for estimating tea-plantation N2O emissions and quantitative support for sustainable nitrogen management and targeted greenhouse gas mitigation strategies in tea production systems.

Xiao-Ting Jie, Xin Liu, Jian-Fei Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.