A realistic performance boundary is outlined for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models.
Abstract
Machine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We construct a reproducible benchmark that fuses 936 CatApp DFT adsorption energies with bulk descriptors from the Materials Project for H*, O*, and OH* on metal and alloy surfaces. We compare random K-fold cross-validation with GroupKFold grouped by parsed formula, the latter mimicking the realistic task of predicting adsorption on entirely new catalyst compositions. Under formula-grouped evaluation, random CV materially overestimates apparent generalization performance, with the largest and most robust effects for H* and OH* (protocol-inflation gaps up to approximately 0.8). The H* and OH* results are based on only 20 and 31 unique formulas, so their GroupKFold Spearman point estimates should be read as directional evidence rather than quantitative estimates. O* shows a smaller and statistically fragile protocol-inflation signal and, even where composition-plus-bulk features improve Random Forest and Ridge, the usable signal is best described as a very coarse pre-filter within a limited domain. Bulk descriptors are adsorbate-dependent: they improve O* prediction for Random Forest and Ridge, but degrade H* and OH*—a qualitative, directional observation given the small formula counts—whose binding is poorly captured by bulk crystal descriptors, consistent with the established view that it is governed by surface-localized electronic structure. These results outline a realistic performance boundary for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models. We therefore recommend that bulk-to-surface adsorption-energy benchmarks report formula-grouped cross-validation alongside random cross-validation as a more robust and transparent practice.
Despite the rapid adoption of machine learning (ML) in materials discovery, its application to water contaminant adsorption remains fundamentally constrained. In this study, we systematically evaluated whether current peer-reviewed literature on magnetite nanoparticle (MNPs) adsorbents contains sufficient descriptor...
A. I. Yunus, Jacob Song, Samuel Darko et al.· ACS ES&T Water· 0 citations
This work presents a machine learning-symbolic regression strategy to develop a physically interpretable formula for predicting low pressure CO2 adsorption capacity in hypothetical metal-organic frameworks (hMOFs), and proposes a physics-guided expression that enables efficient prediction and provides clearer insight i...
Yimin Shao, Sheng-Ling Ma, Sheng-Hong Ju et al.· 0 citations
Identifying descriptors associated with gaseous arsenic adsorption by metal oxides remains challenging because literature data are heterogeneous and incomplete. A database of 280 experimental records and 20 descriptors from 17 studies was compiled to predict adsorption capacity and interpret descriptor-performance rela...
Yanhong Zhu, Qi Liu, Shuang-Chun Wen et al.· Journal of Environmental Man...· 0 citations
A global active learning framework is demonstrated to map these landscapes efficiently by coupling genetic algorithms with deep neural networks trained on density functional theory data, which provides a scalable and resource-efficient strategy for high-throughput materials discovery in applications such as hydrogen st...
J. von der Heyde, Walter Malone, A. Kara· Journal of Physical Chemistr...· 0 citations
Metal–organic frameworks (MOFs) are promising materials for adsorption and separation, but accurately predicting water adsorption remains a major challenge in molecular simulation. Classical force fields, such as UFF, often fail to capture the strong, directional hydrogen-bonding interactions between water and MOFs, wh...
Yu-Tao Li, Xiao-Qi Zhang, Xin Jin et al.· Journal of Chemical Theory a...· 0 citations
An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.
D. Pundhir, Ashok Kumar· Applied Physics A· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.