Skip to content
Open access

Reliable Machine Learning Screening of Adsorption Energies Is Better Assessed with Formula-Grouped Cross-Validation

Aug 2026 · Catalysts · 0 citations

TL;DR

A realistic performance boundary is outlined for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models.

Abstract

Machine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We construct a reproducible benchmark that fuses 936 CatApp DFT adsorption energies with bulk descriptors from the Materials Project for H*, O*, and OH* on metal and alloy surfaces. We compare random K-fold cross-validation with GroupKFold grouped by parsed formula, the latter mimicking the realistic task of predicting adsorption on entirely new catalyst compositions. Under formula-grouped evaluation, random CV materially overestimates apparent generalization performance, with the largest and most robust effects for H* and OH* (protocol-inflation gaps up to approximately 0.8). The H* and OH* results are based on only 20 and 31 unique formulas, so their GroupKFold Spearman point estimates should be read as directional evidence rather than quantitative estimates. O* shows a smaller and statistically fragile protocol-inflation signal and, even where composition-plus-bulk features improve Random Forest and Ridge, the usable signal is best described as a very coarse pre-filter within a limited domain. Bulk descriptors are adsorbate-dependent: they improve O* prediction for Random Forest and Ridge, but degrade H* and OH*—a qualitative, directional observation given the small formula counts—whose binding is poorly captured by bulk crystal descriptors, consistent with the established view that it is governed by surface-localized electronic structure. These results outline a realistic performance boundary for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models. We therefore recommend that bulk-to-surface adsorption-energy benchmarks report formula-grouped cross-validation alongside random cross-validation as a more robust and transparent practice.

Read PDF

Similar papers

Review Open access Aug 2026

Machine Learning for Magnetite Nanoparticles Contaminant Adsorption under Uncertainty: Revealing Data Quality Limits

Despite the rapid adoption of machine learning (ML) in materials discovery, its application to water contaminant adsorption remains fundamentally constrained. In this study, we systematically evaluated whether current peer-reviewed literature on magnetite nanoparticle (MNPs) adsorbents contains sufficient descriptor...

A. I. Yunus, Jacob Song, Samuel Darko et al. · 0 citations
Preprint Aug 2026

Discovering Physically Interpretable Mathematical Expression for Predicting CO2 Adsorption in Metal-Organic Frameworks via Machine Learning-Symbolic Regression

This work presents a machine learning-symbolic regression strategy to develop a physically interpretable formula for predicting low pressure CO2 adsorption capacity in hypothetical metal-organic frameworks (hMOFs), and proposes a physics-guided expression that enables efficient prediction and provides clearer insight i...

Yimin Shao, Sheng-Ling Ma, Sheng-Hong Ju et al. · 0 citations
Aug 2026

Interpretable machine learning for predicting gaseous arsenic adsorption by metal oxides and identifying influential descriptors.

Identifying descriptors associated with gaseous arsenic adsorption by metal oxides remains challenging because literature data are heterogeneous and incomplete. A database of 280 experimental records and 20 descriptors from 17 studies was compiled to predict adsorption capacity and interpret descriptor-performance rela...

Yanhong Zhu, Qi Liu, Shuang-Chun Wen et al. · 0 citations
Aug 2026

Active Learning for Optimizing Adsorption Energy Predictions in Large Chemical Spaces.

A global active learning framework is demonstrated to map these landscapes efficiently by coupling genetic algorithms with deep neural networks trained on density functional theory data, which provides a scalable and resource-efficient strategy for high-throughput materials discovery in applications such as hydrogen st...

J. von der Heyde, Walter Malone, A. Kara · 0 citations
Open access Aug 2026

Generalized Machine Learning Potentials for Predicting Low-Pressure Water Adsorption in Flexible Al-Based Metal–Organic Frameworks

Metal–organic frameworks (MOFs) are promising materials for adsorption and separation, but accurately predicting water adsorption remains a major challenge in molecular simulation. Classical force fields, such as UFF, often fail to capture the strong, directional hydrogen-bonding interactions between water and MOFs, wh...

Yu-Tao Li, Xiao-Qi Zhang, Xin Jin et al. · 0 citations
Jul 2026

Predicting dielectric constants of crystalline materials using explainable machine learning and composition-aware feature engineering

An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.

D. Pundhir, Ashok Kumar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.