Skip to content
Preprint

Computed materials proposals depart from the structural memory of experimental discovery

Jun 2026 · 0 citations
Physics

TL;DR

This work embeds 167,500 Inorganic Crystal Structure Database entries in a continuous structural-similarity space, partition it into graph communities, and replay them in time to define a historical synthesizability prior for triaging computed materials.

Abstract

Generative AI and high-throughput DFT pipelines propose millions of inorganic crystal structures, but lack a calibrated reference frame against experimentally realized chemistry. Here we embed 167,500 Inorganic Crystal Structure Database entries in a continuous structural-similarity space, partition it into graph communities, and replay them in time. Experimental discovery shows strong structural memory: 82.9% of new formulas enter pre-existing communities; new-community formation falls from 40.2% (1930s) to 2.6% (2010s). The communities are chemically meaningful, positively identifying nine textbook field-defining renaissances, including cuprates, colossal-magnetoresistance manganites, MAX phases, and Li-ion battery cathodes. Projecting GNoME, MatterGen-public, Materials Project, JARVIS-DFT, and Alexandria-PBE into frozen historical maps yields a cutoff-robust ordering: held-out ICSD>MatterGen>{GNoME ~ MP-theoretical}>JARVIS>Alexandria. Structural departure from experimental basins is not specific to generative AI but general across the tested computed sets. Combining structural proximity with reduced-formula precedent defines a historical synthesizability prior for triaging computed materials.

View source

Similar papers

Preprint Jun 2026

Adaptive fine-tuning of foundation models for crystal structure prediction: Discovery of high-pressure phases in the CaFeNi system

The prediction of crystal structures is a key challenge in chemistry and materials science, but evolutionary crystal structure prediction (CSP) remains computationally expensive because it relies on repeated \textit{ab initio} relaxations and energy ranking. Machine learning interatomic potentials (MLIPs) can accelerate CSP, yet their use is limited by the need for large training sets and by the difficulty of choosing which candidate structures should be labeled by density functional theory (DFT). Here we introduce a self-consistent, foundation-model-assisted CSP workflow that combines evolutionary search with adaptive data selection and fine-tuning. Starting from a pretrained MLIP, the algorithm rapidly explores configuration space while iteratively selecting compact, representative, and physically relevant subsets of structures for DFT labeling, thereby reducing redundant calculations and improving a system-specific potential. We apply the method to the chemically complex Ca--Fe--Ni ternary system. The workflow reproduces the known low-pressure convex hull and enables efficient high-pressure exploration. It predicts a previously unreported compound, Ca$_6$FeNi, which becomes thermodynamically stable above 100~GPa. These results show that foundation-model-based, data-efficient CSP can greatly reduce computational cost while preserving accuracy and enabling the discovery of new materials in complex multicomponent systems.

N. Chtchelkatchev, M. Magnitskaya, R. Ryltsev · 0 citations
Preprint Jul 2026

Correcting DFT formation energies towards experimental accuracy using foundational MLIPs and latent-feature delta-learning

This work demonstrates how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation.

Timo Reents, Marnik Bercx, Giovanni Pizzi · 0 citations
Preprint Jul 2026

Predicting Novel Stable Materials for Experimental Synthesis

Machine-learning-accelerated materials discovery has yielded large numbers of computationally stable compounds, yet many remain experimentally unrealized, underscoring a persistent gap between prediction and synthesis. Here, we introduce a hierarchical screening framework that combines PBE-based thermodynamic stability, efficient dynamical-stability screening enabled by universal machine-learning interatomic potentials, and SCAN-based thermodynamic refinement. Applying this protocol to the 894 stable materials previously reported in Sci. Data 9, 302 (2022), we first curate 603 unique structures, of which only 298 remain thermodynamically stable on the complete PBE phase diagrams, demonstrating the critical role of competing phases in stability assessment. Dynamical screening then identifies 166 materials stable under both harmonic-phonon and finite-temperature molecular dynamics criteria, and SCAN phase diagrams further narrow this set to 109. Finally, by combining decomposition enthalpy with chemical-space completeness, we prioritize 25 candidates as high-confidence targets for experimental synthesis. This work provides a practical protocol for translating stability predictions into experimentally actionable synthesis targets, closing a key gap in machine-learning-driven materials discovery.

Yuqi An, Sihong Zhu, Joseph H. Montoya et al. · 0 citations
Jul 2026

Active-Learning Discovery of Superionic Compositions Using High-Throughput EIS and Structure-Aware Descriptors

We introduce an active-learning framework that closes the loop between high-throughput EIS measurements and structure-aware composition descriptors to discover superionic candidates under realistic processing constraints. Starting from a small seed set, Gaussian-process and tree-based models propose batched experiments that maximize information gain on conductivity and activation energy while enforcing uncertainty-aware Kramers–Kronig quality gates. Descriptor families integrate interpretable features: ionic radius mismatch, framework softness, site connectivity from simple graph-derived motifs, and processing proxies (grain size from Scherrer, porosity, interphase penalty terms). We demonstrate rapid convergence to high-conductivity regions in multi-component chalcogenide and halide spaces using the automated multi-site EIS workflow described separately. Across three material spaces, the approach reduces experiments ~3× versus grid sampling while yielding candidates with improved conductivity at moderate temperatures and stable impedance upon cycling. We release a lightweight, reproducible stack (metadata schema, analysis notebooks, and synthetic datasets) to encourage community benchmarking without proprietary infrastructure. The result is a pragmatic path to self-driving electrolyte discovery that prioritizes experimental tractability and interpretability—features that matter for industrial translation and cross-lab reproducibility. Keywords: active learning; Bayesian optimization; EIS QC; interpretable descriptors; high-throughput screening; solid electrolytes

Progna Banerjee · 0 citations
Preprint Jul 2026

Stoichiometric cluster learning for few-shot property prediction of multi-ionic integrated energetic materials

It is shown how pretrained machine-learned interatomic potentials (MLIPs) can bypass full crystal-structure prediction and support pre-synthesis screening from stoichiometric ionic clusters using multi-ionic integrated explosives (MIXs) as a synthesis-facing example.

Ming-Yu Guo, W. Zou, Yu Shang et al. · 0 citations