Skip to content
Review Open access

Illusions of Accuracy: Widespread Data Leakage Distorts Machine Learning Results in Biochar Research

Aug 2026 · GCB Bioenergy · Vol 18 · 0 citations · 55 references

TL;DR

To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks.

Abstract

Machine learning (ML) is increasingly used in biochar research to predict yield, material properties, and adsorption performance. Yet, this review reveals that much of the reported success in ML‐driven biochar studies may be overstated due to data leakage, the inadvertent transfer of information between training and testing datasets. We systematically examined 64 of the most‐cited studies published between 2015 and 2025 and found that all studies exhibited potential leakage, commonly arising from hierarchical or grouped data structures, multi‐source compilations, or temporal and concentration dependencies. Nearly all studies (98%) employed random point‐level data splits for training and testing, disregarding dependencies among samples, while only one adopted a group‐based split approach. Re‐analysis of five representative datasets showed that random point‐level splitting can inflate predictive accuracy by up to 86%, often yielding test‐set R2 (coefficient of determination) values exceeding 0.90 even when model generalisation fails. Leakage risks were particularly severe in adsorption studies and in those using very small datasets (< 35 samples), where overfitting further distorted model performance. To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks. Future progress requires automated leakage‐detection tools, benchmark datasets with explicit metadata on dependencies, and hybrid models that integrate physical constraints. Addressing data leakage systematically is essential to restore credibility, enhance reproducibility, and realise the full potential of ML as a reliable tool for environmental research.

Read PDF

Similar papers

Open access Sep 2026

Leakage-aware machine learning for data-driven performance prediction of metal–organic framework systems

Metal–organic frameworks (MOFs) are highly tunable porous materials whose performance is governed by complex interactions among structural, chemical, material, and operating variables. This study develops a leakage-aware, data-driven framework for predicting two distinct MOF performance endpoints: loading capacity an...

N. Abu-Hamdeh, M. Ajour, A. Milyani et al. · 0 citations
Review Open access Aug 2026

Machine Learning-Driven Advances in Hydrogen Embrittlement of Steels: A Comprehensive Review

This paper systematically reviews and critically evaluates AI/ML applications for HE in steels through a structured analysis of all studies published between 2010 and 2026, and establishes a practical framework for selecting appropriate AI/ML approaches according to dataset characteristics and engineering objectives.

A. G. Talkhan, Fadwa T. Eljack, Seckin Karagoz · 0 citations
Review Open access Aug 2026

Machine Learning for Magnetite Nanoparticles Contaminant Adsorption under Uncertainty: Revealing Data Quality Limits

Despite the rapid adoption of machine learning (ML) in materials discovery, its application to water contaminant adsorption remains fundamentally constrained. In this study, we systematically evaluated whether current peer-reviewed literature on magnetite nanoparticle (MNPs) adsorbents contains sufficient descriptor...

A. I. Yunus, Jacob Song, Samuel Darko et al. · 0 citations
Open access Aug 2026

Machine Learning for Alkali-Activated Concrete: Feature Attribution, Strength–Carbon Relationships, and the Limits of Out-of-Campaign Generalisation

Machine learning (ML) models for alkali-activated concrete (AAC) are almost universally evaluated with random train–test splits, yet the literature-compiled datasets are strongly clustered by source study, and the reliability of such evaluations has rarely been quantified. The novelty of this study is a systematic quan...

F. Pacheco-Torgal, Saqib Iqbal · 0 citations
Sep 2026

Temporal Data Leakage Inflates Machine Learning Performance in Compostable PBAT/PLA Polymer Biodegradation Modeling

This study quantifies how evaluation protocol choice, not model architecture, governs apparent machine learning (ML) performance in poly(butylene adipate-co-terephthalate)/polylactic acid (PBAT/PLA) composite biodegradation modeling. Seven regression architectures spanning linear, regularized, kernel-based, ensemble,...

Jun-Tong Zhang, Su-Wan Chen, Zi-Yi Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.