Skip to content
Open access

Machine learning of electronic structure and atomistic properties from the external potential.

Feb 2026 · Journal of Chemical Physics · Vol 165 3 · 2 citations · 75 references
Physics Medicine

TL;DR

This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.

Abstract

Electronic structure calculations remain a major bottleneck in atomistic simulations and, not surprisingly, have attracted significant attention in machine learning (ML). Most existing approaches learn a direct map from molecular geometries, typically represented as graphs or encoded local environments, to molecular properties or use ML as a surrogate for electronic structure theory by targeting quantities, such as Fock or density matrices expressed in an atomic orbital (AO) basis. Inspired by the Hohenberg-Kohn theorem, in this work, we propose an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input. From this operator, we construct hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors. At the same time, the matrix-valued nature of the external potential provides a natural connection to equivariant message-passing neural networks. In particular, we show that successive products of the external potential provide a scalable route to equivariant message passing and enable an efficient description of nonlocal effects. We demonstrate that this approach can be used to model molecular properties, such as energies and dipole moments, from the external potential or to learn effective operator-to-operator maps, including mappings to the Fock matrix from which multiple molecular observables can be simultaneously derived.

Read PDF

Similar papers

Preprint Jul 2026

Rem3Di: Learning smooth, chiral 3D molecular descriptors from atomistic foundation models

Rem3Di is introduced, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening and provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.

Steffen Wedig, Felix Burton, Rokas Elijošius et al. · 0 citations
Preprint Jul 2026

Extracting Atomic Environments for Machine Learning Interatomic Potentials

A notably simple procedure, a method the authors refer to as deletions, yields superior performance over an array of alternative extraction methods for extracting atomic environments from large, bulk configurations and embedding them into smaller configurations suitable for DFT calculations with periodic boundary conditions.

Jared Stimac, Fei Zhou, Kyle Bushick et al. · 0 citations
Open access Sep 2026

The use of machine learning interatomic potentials for the verification of experimental molecular crystal structures.

A correctly solved crystal structure should agree with the experimental data, and its geometry should correspond to a local minimum on the potential energy surface (PES). The idea of verifying crystal structure solutions by comparing them with their geometry-optimized versions was introduced 15 years ago. Recent developments in machine learning interatomic potentials (MLIPs) have made it possible to replace computationally expensive density functional theory (DFT) calculations with AI/neural-network-based alternatives. MLIPs can reach DFT-comparable precision with a substantial gain in speed. We selected one promising MLIP, Universal Models for Atoms, trained on the Open Molecular Crystals 2025 dataset, and processed a prefiltered subset of 216 919 structures from the Cambridge Structural Database. Due to the limitations of the MLIP available when this study commenced, ionic compounds, salts and metal-containing structures were excluded. The current methodology cannot process disordered structures, and available computational resources limit the maximum unit-cell volume that can be treated to 4000 Å3. All structures in the dataset were geometry optimized using the MLIP, and similarity descriptors were calculated to quantify the differences between the original and optimized structures. Automatic analysis was followed by the manual identification of issues indicated by the descriptors' values. We detected anomalies in experimental structures that had already passed all prior validation, as well as limitations in the reliability of the MLIP PES calculations. For 1867 crystal structures, bond-pattern change was observed, while 3331 structures showed a root-mean-square Cartesian displacement greater than 0.25 Å. Future improvements to the methodology and extension to systems not covered by this study are discussed.

M. Hušák, F. Fňukal, J. Čejka · 0 citations
Preprint Jul 2026

Machine Learning Materials Properties by Encoding Orbital-Projected Density of States

Graph neural networks have become the dominant machine-learning architecture for predicting materials properties from crystal structures. Yet the initialization of atomic node features has received comparatively little attention, and conventional approaches rely on static elemental descriptors that carry no information about the quantum-mechanical electronic environment of each atom in its crystalline host. Here we show that augmenting atomic node representations with site-projected orbital density of states (pDOS) fingerprints, computed directly from density functional theory calculations, yields systematic and substantial improvements in predictive performance.These representations are fused with Pettifor elemental embeddings at each atomic site before message passing. For the superconducting critical temperature $T_c$ and the optical dielectric constant $\epsilon_{\infty}$,the pDOS augmentation reduces prediction errors by 22.9% and 27.9%, respectively, relative to the elemental-descriptor baseline. These improvements are comparable to those achieved by doubling the training-set size. The gains are, however, contingent on training-set size. For the magnetic exchange energies of Heusler compounds, a substantially smaller dataset, the improvement is reduced,indicating that pDOS augmentation is most effective when the training data exceeds the length of the pDOS feature vector. We introduce an interpretable spectral attention-gating mechanism that reveals that the model autonomously learns to prioritize the orbital channels and energy windows most physically relevant to each target property. These results establish pDOS-augmented graph nodes as a broadly applicable strategy for infusing first-principles electronic-structure knowledge into graph networks, opening a practical route to high-accuracy property prediction in data-scarce regimes.

Paulo R. Pires, Pierre-Paul De Breuck, Mauro Fava et al. · 0 citations
Preprint Jul 2026

Reconstructing local environments from concise atomistic representations

Symmetry-based representations of local atomic structure, such as the power spectrum or bispectrum, are routinely used to characterize the structural diversity of datasets and as input features for atomistic machine learning. Although these descriptors systematically incorporate increasingly complex geometric correlations, it remains unclear if a given feature can be mapped back to a discrete point cloud, whether such a reconstruction is unique, and how changes in the descriptor are reflected in the underlying atomic geometry. The choice and discretization of the radial and angular bases, as well as the high dimensionality of the resulting feature vectors -- which may contain hundreds or thousands of components -- make this interpretation even more challenging. In this work, we investigate the inverse problem of recovering atomic structures from local invariant descriptors. We show that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features. Even representations that are formally incomplete or locally ill-conditioned can be inverted to accurate geometric reconstructions of atomic environments across molecular and material datasets. Our reconstruction framework provides a general algorithmic means of identifying approximate degeneracies of invariant descriptors and recovering distinct atomic environments that cannot be distinguished by a given representation. Finally, by reconstructing atomic configurations from descriptors, we examine how perturbations in invariant descriptors of different correlation orders translate into structural distortions.

Jigyasa Nigam, T. Phung, Ameya Daigavane et al. · 1 citation
Preprint Jul 2026

MANDALA: An E(3)-Equivariant Graph Neural Network Framework for Learning Electronic-Structure Operators with Observable Guidance

Electronic-structure calculations based on Kohn-Sham density functional theory remain indispensable in computational materials science and chemistry. Their computational cost, however, limits accessible system sizes and simulation times. At the same time, conventional machine-learning interatomic potentials (MLIPs), which are becoming the workhorse of large-scale materials modeling, usually target only energies and forces. They therefore leave out the quantum-operator-level information required to reconstruct band structures, densities of states, spatial charge distributions, and other electronic observables. \texttt{Mandala} fills this methodological gap. It is a modular software framework for learning block-sparse electronic-structure matrices with E(3)-equivariant graph neural networks. The framework is built around a unified representation of atom-resolved Hamiltonian, overlap, and density matrices, together with reusable abstractions for basis conversion, sparse block handling, irreducible representation mapping, graph construction, model definition, and training. This design allows \texttt{Mandala} to support heterogeneous chemical compositions, a wide range of neural architecture variants within one workflow, and multiple electronic-structure backends. \texttt{Mandala} evaluates selected observables directly from the predicted operators, including band energy, electron count, density of states, and band structure. This connects electronic-structure learning and observable-guided modeling while retaining a representation tied to quantum-mechanical operators rather than only scalar or vector targets as in MLIPs. In this form, \texttt{Mandala} is intended to complement atomistic interatomic potential workflows by resolving electronic structure and operator-derived observables within one scalable implementation.

B. Brzoza, Wiktoria Szopa, Z. Elabid et al. · 0 citations