Aug 2026· GCB Bioenergy· Vol 18· 0 citations· 55 references
TL;DR
To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks.
Abstract
Machine learning (ML) is increasingly used in biochar research to predict yield, material properties, and adsorption performance. Yet, this review reveals that much of the reported success in ML‐driven biochar studies may be overstated due to data leakage, the inadvertent transfer of information between training and testing datasets. We systematically examined 64 of the most‐cited studies published between 2015 and 2025 and found that all studies exhibited potential leakage, commonly arising from hierarchical or grouped data structures, multi‐source compilations, or temporal and concentration dependencies. Nearly all studies (98%) employed random point‐level data splits for training and testing, disregarding dependencies among samples, while only one adopted a group‐based split approach. Re‐analysis of five representative datasets showed that random point‐level splitting can inflate predictive accuracy by up to 86%, often yielding test‐set R2 (coefficient of determination) values exceeding 0.90 even when model generalisation fails. Leakage risks were particularly severe in adsorption studies and in those using very small datasets (< 35 samples), where overfitting further distorted model performance. To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks. Future progress requires automated leakage‐detection tools, benchmark datasets with explicit metadata on dependencies, and hybrid models that integrate physical constraints. Addressing data leakage systematically is essential to restore credibility, enhance reproducibility, and realise the full potential of ML as a reliable tool for environmental research.
A validation ladder that reports row-wise interpolation, specified-group transfer, and independent external performance makes the decision boundary of an aqueous-removal model explicit.
Siyuan Jiang, Ying Yang, Yao-Ru Mao et al.· Journal of Hazardous Materia...· 0 citations
Metal–organic frameworks (MOFs) are highly tunable porous materials whose performance is governed by complex interactions among structural, chemical, material, and operating variables. This study develops a leakage-aware, data-driven framework for predicting two distinct MOF performance endpoints: loading capacity an...
N. Abu-Hamdeh, M. Ajour, A. Milyani et al.· Frontiers in Medicine· 0 citations
This paper systematically reviews and critically evaluates AI/ML applications for HE in steels through a structured analysis of all studies published between 2010 and 2026, and establishes a practical framework for selecting appropriate AI/ML approaches according to dataset characteristics and engineering objectives.
A. G. Talkhan, Fadwa T. Eljack, Seckin Karagoz· Hydrogen· 0 citations
Despite the rapid adoption of machine learning (ML) in materials discovery, its application to water contaminant adsorption remains fundamentally constrained. In this study, we systematically evaluated whether current peer-reviewed literature on magnetite nanoparticle (MNPs) adsorbents contains sufficient descriptor...
A. I. Yunus, Jacob Song, Samuel Darko et al.· ACS ES&T Water· 0 citations
Machine learning (ML) models for alkali-activated concrete (AAC) are almost universally evaluated with random train–test splits, yet the literature-compiled datasets are strongly clustered by source study, and the reliability of such evaluations has rarely been quantified. The novelty of this study is a systematic quan...
F. Pacheco-Torgal, Saqib Iqbal· Construction Materials· 0 citations