The prediction of crystal structures is a key challenge in chemistry and materials science, but evolutionary crystal structure prediction (CSP) remains computationally expensive because it relies on repeated \textit{ab initio} relaxations and energy ranking. Machine learning interatomic potentials (MLIPs) can accelerate CSP, yet their use is limited by the need for large training sets and by the difficulty of choosing which candidate structures should be labeled by density functional theory (DFT). Here we introduce a self-consistent, foundation-model-assisted CSP workflow that combines evolutionary search with adaptive data selection and fine-tuning. Starting from a pretrained MLIP, the algorithm rapidly explores configuration space while iteratively selecting compact, representative, and physically relevant subsets of structures for DFT labeling, thereby reducing redundant calculations and improving a system-specific potential. We apply the method to the chemically complex Ca--Fe--Ni ternary system. The workflow reproduces the known low-pressure convex hull and enables efficient high-pressure exploration. It predicts a previously unreported compound, Ca$_6$FeNi, which becomes thermodynamically stable above 100~GPa. These results show that foundation-model-based, data-efficient CSP can greatly reduce computational cost while preserving accuracy and enabling the discovery of new materials in complex multicomponent systems.
This work demonstrates how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation.
Timo Reents, Marnik Bercx, Giovanni Pizzi· 0 citations
A correctly solved crystal structure should agree with the experimental data, and its geometry should correspond to a local minimum on the potential energy surface (PES). The idea of verifying crystal structure solutions by comparing them with their geometry-optimized versions was introduced 15 years ago. Recent developments in machine learning interatomic potentials (MLIPs) have made it possible to replace computationally expensive density functional theory (DFT) calculations with AI/neural-network-based alternatives. MLIPs can reach DFT-comparable precision with a substantial gain in speed. We selected one promising MLIP, Universal Models for Atoms, trained on the Open Molecular Crystals 2025 dataset, and processed a prefiltered subset of 216 919 structures from the Cambridge Structural Database. Due to the limitations of the MLIP available when this study commenced, ionic compounds, salts and metal-containing structures were excluded. The current methodology cannot process disordered structures, and available computational resources limit the maximum unit-cell volume that can be treated to 4000 Å3. All structures in the dataset were geometry optimized using the MLIP, and similarity descriptors were calculated to quantify the differences between the original and optimized structures. Automatic analysis was followed by the manual identification of issues indicated by the descriptors' values. We detected anomalies in experimental structures that had already passed all prior validation, as well as limitations in the reliability of the MLIP PES calculations. For 1867 crystal structures, bond-pattern change was observed, while 3331 structures showed a root-mean-square Cartesian displacement greater than 0.25 Å. Future improvements to the methodology and extension to systems not covered by this study are discussed.
This work not only establishes a pioneering paradigm for interpretable ML-driven force field refinement but also provides the first feature engineering solution incorporating chemical, physical, and structural information specifically designed for the machine learning of energetic molecular crystals.
Qi He, Pengju Wang, Xudong He et al.· Molecules· 0 citations
Molecular crystal structure prediction (CSP) is important in pharmaceuticals, agrochemicals, and organic electronics, where subtle differences in molecular conformation and packing can strongly affect material properties. We present Packora, a flow-based generative model for molecular CSP that jointly predicts atomic coordinates and the lattice from molecular graphs. Packora supports multi-component and organometallic crystals and can condition on any subset of molecular conformers, stereochemical labels, and space-group information within a single model. Inspired by the CCDC CSP blind test, we evaluate generation and ranking separately, using generation to isolate generator quality and ranking to measure end-to-end performance under a common relaxation and ranking pipeline. We also systematically study architecture, training, conditioning, inference, and scaling, identifying an effective design based on cacheable pairwise reasoning, training objective and numerical solver choices, conditioning dropout, and balanced scaling of pairwise and single representations. Packora outperforms the baselines on both structure generation and ranking benchmarks, achieving the best matched-budget coverage across all six generation benchmarks, as well as higher experimental-form recovery, lower experimental-form ranks, and faster convergence in ranking.
Nayoung Kim, Kiyoung Seong, Sungsoo Ahn· 0 citations
The discovery and design of novel transition metal complexes for specific applications heavily rely on computational high-throughput screenings to identify promising candidates for experimental validation. However, traditional computational approaches, such as density functional theory, are often too computationally demanding to be applied on a large scale. Machine learning methods offer a promising alternative due to their excellent computational efficiency, but their accuracy and high data requirements remain major challenges for their effective implementation. To address these issues, we herein present an adaptation of the Δ-ML strategy for quantum property prediction of transition metal complexes. We combine GFN2-xTB geometry optimizations and density functional theory single-point calculations in order to obtain low-fidelity approximations and generate featurized graph representations that serve as input to a graph neural network architecture. The high-fidelity targets originate from the tmQMg dataset and include the electronic and dispersion energies, HOMO-LUMO gap and dipole moment at the PBE0-D3BJ/def2-TZVP level as well as the polarizability at the PBE-D3BJ/def2-SVP level. Compared to a conventional benchmark approach, the proposed method consistently achieves higher accuracy in the prediction of high-fidelity targets, while demonstrating improved data efficiency and out-of-domain transferability. We furthermore show, how the use of cheaper low-fidelity methods leads to significant reductions in computational cost at minor losses in predictive performance. Overall, these results highlight the potential of Δ-ML for materials discovery in transition metal chemistry, which requires high predictive accuracy despite often times limited availability of training data.
Hannes Kneiding, David Balcells· Chemistry· 0 citations