Reliable Machine Learning Screening of Adsorption Energies Is Better Assessed with Formula-Grouped Cross-Validation
A realistic performance boundary is outlined for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models.