Graph neural networks have become the dominant machine-learning architecture for predicting materials properties from crystal structures. Yet the initialization of atomic node features has received comparatively little attention, and conventional approaches rely on static elemental descriptors that carry no information about the quantum-mechanical electronic environment of each atom in its crystalline host. Here we show that augmenting atomic node representations with site-projected orbital density of states (pDOS) fingerprints, computed directly from density functional theory calculations, yields systematic and substantial improvements in predictive performance.These representations are fused with Pettifor elemental embeddings at each atomic site before message passing. For the superconducting critical temperature $T_c$ and the optical dielectric constant $\epsilon_{\infty}$,the pDOS augmentation reduces prediction errors by 22.9% and 27.9%, respectively, relative to the elemental-descriptor baseline. These improvements are comparable to those achieved by doubling the training-set size. The gains are, however, contingent on training-set size. For the magnetic exchange energies of Heusler compounds, a substantially smaller dataset, the improvement is reduced,indicating that pDOS augmentation is most effective when the training data exceeds the length of the pDOS feature vector. We introduce an interpretable spectral attention-gating mechanism that reveals that the model autonomously learns to prioritize the orbital channels and energy windows most physically relevant to each target property. These results establish pDOS-augmented graph nodes as a broadly applicable strategy for infusing first-principles electronic-structure knowledge into graph networks, opening a practical route to high-accuracy property prediction in data-scarce regimes.
This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.
Jigyasa Nigam, T. Smidt, G. Dusson· Journal of Chemical Physics· 2 citations
Results indicate that the topology-aligned inductive bias is the active ingredient driving parameter efficiency at QM9 scale, with implications for matched-baseline benchmarking in quantum machine learning.
Foundation machine-learning interatomic potentials (MLIPs) enable atomistic simulations at substantially lower computational cost than first-principles methods, but their reliability across structural geometries remains insufficiently understood. Here, we construct a density-functional-theory dataset of ZrO2 configurations spanning bulk, slab, particle, neck, and atomically thin wire environments motivated by an experimentally observed ZrO2 desintering process involving neck thinning and atomic wire formation. We first benchmark 26 pretrained MLIPs and observe pronounced geometry-dependent degradation in zero-shot predictions. Without any training, after only reference-energy alignment, the best zero-shot model (ORB-V3) reaches energy and force root-mean-square errors of 6 meV/atom and 197.3 meV/{\AA}, respectively, with the largest force errors in neck and wire configurations. We then compare zero-shot inference, fine-tuning, and training from scratch strategies. Fine-tuning yields lower energy and force errors than training from scratch, while both require comparable wall-clock time. Geometry-specific fine-tuning improves in-domain accuracy but frequently produces negative transfer to other structural classes, whereas mixed-geometry fine-tuning reduces cross-geometry errors. Evaluations of elastic and vibrational properties, surface energies, and neck dynamics further show that rankings based on average energy and force errors do not universally predict property-level behavior. These results demonstrate that geometry-diverse target data and independent physical validations are necessary when adapting foundation MLIPs to low-coordination (ionic) nanostructures.
P. Zanineli, B. Focassio, G. R. Schleder· 0 citations
Transformer Atomic Cluster Expansion (TRACE) is introduced, an energy-conserving architecture that combines atomic cluster expansion density correlations with local multihead cross-attention that captures multi-species crystallization, liquid structures, phase diagrams, and chemical reactivity.
The prediction of the structural stability of octet $AB$-type binary compounds is a classical materials informatics problem. The challenge is to capture the relative stability of 4-fold coordinated atoms in zincblende ($\beta$-ZnS) structure and 6-fold coordinated atoms in rocksalt (NaCl) structure, modulated by charge transfer and atomic-size differences. Previous structure maps and machine-learning approaches used atomic features such as valence-electron count, ionization potential and atomic radii, using either physical intuition or symbolic regression. Here, we demonstrate that explicitly incorporating the domain knowledge of the interatomic bonds can significantly and systematically improve the prediction of $\beta$-ZnS/NaCl stability. We encode this bonding information through a coarse-grained representation of the local electronic structure obtained by a recursive solution of a tight-binding bond model. The underlying pairwise Hamiltonians are taken from downfolded eigenstates of density-functional theory calculations for diatomic molecules and thereby include domain knowledge of the bond between specific $A-B$ pairs. The benefit of this description is demonstrated with an ensemble of independently trained Kernel Ridge or symbolic regression models combined with sequential feature selection. The obtained models are compared to a previous symbolic-regression model using the same set of \emph{ab initio} calculations for octet binaries as training data. We find a significant improvement in the prediction of the formation energy difference of $AB$ compounds as compared to previous works and demonstrate that an increasing amount of bond-informed recursion features improves the predictive accuracy.
Rohan D. Kumar, Mariano Forti, A. Naik et al.· 0 citations
Rem3Di is introduced, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening and provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.
Steffen Wedig, Felix Burton, Rokas Elijošius et al.· 0 citations