This work demonstrates how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation.
Abstract
Crystal structure databases curated by high-throughput density functional theory calculations typically serve as the starting point for computational materials discovery efforts. Thermodynamic stability data, such as formation energies and the energy above the convex hull, are important quantities to guide the search for novel materials, enabling filtering for (meta)stable structures. Here, we present the thermodynamic stability of the fully open-source, reproducible, and experimentally focused Materials Cloud three-dimensional crystals database (MC3D). We compare against two other DFT databases, the Open Quantum Materials Database (OQMD) and the Materials Project (MP), as well as against experimental formation enthalpies. We then demonstrate how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level (specifically, we test PET-OMATPES here) can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation. Our results validate and extend the established practice of combining PBEsol geometries with meta-GGA energies to the era of foundational MLIPs. Finally, we train classical machine learning models to further correct the formation energies in a delta-learning framework, where we use the information-rich latent features of the foundational MLIP. These models further reduce the mean absolute error below 50 meV/atom, bringing it down to values comparable with the experimental uncertainty itself. Notably, compared to purely compositional features, the latent features (combined with carefully tuned regularization) simultaneously reduce the prediction error and limit the impact of the learned corrections on the relative phase stability.
A correctly solved crystal structure should agree with the experimental data, and its geometry should correspond to a local minimum on the potential energy surface (PES). The idea of verifying crystal structure solutions by comparing them with their geometry-optimized versions was introduced 15 years ago. Recent developments in machine learning interatomic potentials (MLIPs) have made it possible to replace computationally expensive density functional theory (DFT) calculations with AI/neural-network-based alternatives. MLIPs can reach DFT-comparable precision with a substantial gain in speed. We selected one promising MLIP, Universal Models for Atoms, trained on the Open Molecular Crystals 2025 dataset, and processed a prefiltered subset of 216 919 structures from the Cambridge Structural Database. Due to the limitations of the MLIP available when this study commenced, ionic compounds, salts and metal-containing structures were excluded. The current methodology cannot process disordered structures, and available computational resources limit the maximum unit-cell volume that can be treated to 4000 Å3. All structures in the dataset were geometry optimized using the MLIP, and similarity descriptors were calculated to quantify the differences between the original and optimized structures. Automatic analysis was followed by the manual identification of issues indicated by the descriptors' values. We detected anomalies in experimental structures that had already passed all prior validation, as well as limitations in the reliability of the MLIP PES calculations. For 1867 crystal structures, bond-pattern change was observed, while 3331 structures showed a root-mean-square Cartesian displacement greater than 0.25 Å. Future improvements to the methodology and extension to systems not covered by this study are discussed.
The discovery and design of novel transition metal complexes for specific applications heavily rely on computational high-throughput screenings to identify promising candidates for experimental validation. However, traditional computational approaches, such as density functional theory, are often too computationally demanding to be applied on a large scale. Machine learning methods offer a promising alternative due to their excellent computational efficiency, but their accuracy and high data requirements remain major challenges for their effective implementation. To address these issues, we herein present an adaptation of the Δ-ML strategy for quantum property prediction of transition metal complexes. We combine GFN2-xTB geometry optimizations and density functional theory single-point calculations in order to obtain low-fidelity approximations and generate featurized graph representations that serve as input to a graph neural network architecture. The high-fidelity targets originate from the tmQMg dataset and include the electronic and dispersion energies, HOMO-LUMO gap and dipole moment at the PBE0-D3BJ/def2-TZVP level as well as the polarizability at the PBE-D3BJ/def2-SVP level. Compared to a conventional benchmark approach, the proposed method consistently achieves higher accuracy in the prediction of high-fidelity targets, while demonstrating improved data efficiency and out-of-domain transferability. We furthermore show, how the use of cheaper low-fidelity methods leads to significant reductions in computational cost at minor losses in predictive performance. Overall, these results highlight the potential of Δ-ML for materials discovery in transition metal chemistry, which requires high predictive accuracy despite often times limited availability of training data.
Hannes Kneiding, David Balcells· Chemistry· 0 citations
A machine learning approach is presented that accelerates DFTB simulations by predicting optimal initial atomic charges and demonstrates that ML-predicted initial charges consistently and significantly improve SCC convergence across diverse chemical systems including organic molecules, biomolecules, water clusters, transition metal oxides and solid electrolytes.
Maximilian L. Ach, Karsten Reuter, C. Panosetti· 0 citations
DensIP is introduced, a physics-based model of intermolecular interactions that uses machine-learned electron densities and only four universal parameters that outperforms state-of-the-art general-purpose MLFFs for long-range interactions and can be applied to molecules as large as drug ligands.
Dahvyd Wing, Mihail Bogojeski, Szabolcs Góger et al.· 0 citations
This work introduces MLIP Studio, an open and free platform that brings more than 60 universal MLIPs into a unified interactive interface for molecules and materials, and demonstrates that MLIP-based pre-optimization can reduce subsequent DFT optimization effort by ~33$\times$.
Manas Sharma, Sudeep N. Punnathanam, A. Rajan· 1 citation
Density-functional theory (DFT) has been the workhorse of first-principles calculations for decades, and DFT-derived energies and forces are now widely used to train machine learning models of inter-atomic potentials. However, DFT’s single-particle treatment of exchange-correlation functionals severely limits accuracy for materials with open d- and f-shell elements, and ML models trained on such data inherit this limitation. Dynamical mean-field theory (DMFT) addresses this limitation by explicitly incorporating local electronic correlations, albeit at a significantly higher computational cost. In this work, we develop deep-learning models trained on ab-initio DFT+DMFT calculations to predict electronic self-energies from non-interacting Green’s functions. Using the correlated metal SrVO
3
as a prototype, we show that accurate self-energy predictions can be achieved from small datasets. Through transfer-learning, models pre-trained on SrVO
3
successfully predict the self-energies of CaVO
3
, BaVO
3
and SrNbO
3
, despite differences in composition and electronic structure. Moreover, models pretrained on SrVO
3
and SrNbO
3
can predict self-energy of BaNbO
3
without training on its self-energy. This approach captures temperature variation, extends beyond d
1
perovskites and drastically reduces computational time. These results establish deep-learning as an efficient surrogate for computationally demanding DMFT calculations, enabling rapid prediction of correlation-driven properties, paving the way for a transformative shift in materials theory.