Sep 2026· Journal of Hazardous Materials· Vol 517, pp.
143594
· 0 citations· 28 references
Medicine
TL;DR
A validation ladder that reports row-wise interpolation, specified-group transfer, and independent external performance makes the decision boundary of an aqueous-removal model explicit.
Abstract
Machine-learning models are increasingly used to rank adsorbents and treatment conditions for removing hazardous contaminants from water. A random row split can, however, place measurements from the same material, pollutant, or source study in both training and testing, so a high score may describe interpolation rather than transfer. We surveyed 75 aqueous-removal studies and re-tested five reconstructable adsorption datasets while holding the rows, physically available features, learner, hyperparameters, and seed fixed. Sixty-four studies (85.3%) used validation that did not preserve a reported entity boundary. Replacing shuffled five-fold validation with leave-one-specified-group-out testing reduced pooled R² by 0.108-0.983. One benchmark retained R² = 0.878 under chemical-signature holdout, whereas a uranium-biochar compilation fell from 0.981 to approximately zero under source-study holdout. An independent model-data transfer test on 60 condition means from four adsorbents gave R² = -7.187 after source-study holdout had given 0.510. Because dose was absent from the published feature set and external descriptors and laboratory procedures differed, this external result does not isolate splitting as the cause. A validation ladder that reports row-wise interpolation, specified-group transfer, and independent external performance makes the decision boundary of an aqueous-removal model explicit.
To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks.
H. Jang, Wartini Ng, Song-Chao Chen et al.· GCB Bioenergy· 0 citations
A realistic performance boundary is outlined for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models.
Despite the rapid adoption of machine learning (ML) in materials discovery, its application to water contaminant adsorption remains fundamentally constrained. In this study, we systematically evaluated whether current peer-reviewed literature on magnetite nanoparticle (MNPs) adsorbents contains sufficient descriptor...
A. I. Yunus, Jacob Song, Samuel Darko et al.· ACS ES&T Water· 0 citations
Hydrothermal liquefaction (HTL) distributes wet biomass and residues among bio-oil, char, aqueous product, and gas. Independent phase regressors can produce negative or non-closed allocations, and random validation obscures inter-study heterogeneity. We develop a conservation-informed constrained surrogate treating the...
Machine learning (ML) models for alkali-activated concrete (AAC) are almost universally evaluated with random train–test splits, yet the literature-compiled datasets are strongly clustered by source study, and the reliability of such evaluations has rarely been quantified. The novelty of this study is a systematic quan...
F. Pacheco-Torgal, Saqib Iqbal· Construction Materials· 0 citations
The time-dependent removal performance under irradiation of activated-carbon-supported nickel oxide (AC@NiO) samples containing nominal Cu-loadings of 0%, 3%, 5%, 7.5%, and 10% was examined using experimental concentration data and an explainable machine-learning (ML) framework. Reaction time and nominal Cu-loading wer...
Nesrin Bulut, M. Zontul, Seda Karateke et al.· Catalysts· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.