Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.
A machine learning approach is presented that accelerates DFTB simulations by predicting optimal initial atomic charges and demonstrates that ML-predicted initial charges consistently and significantly improve SCC convergence across diverse chemical systems including organic molecules, biomolecules, water clusters, transition metal oxides and solid electrolytes.
Maximilian L. Ach, Karsten Reuter, C. Panosetti· 0 citations
New methods and workflows to overcome the challenges inherent to automating unrestricted coupled cluster calculations are developed and a transferable MLIP for gas-phase reactions, trained on unrestricted CCSD(T) data is developed.
Alice E. A. Allen, Rui Li, Sakib Matin et al.· Journal of Chemical Theory a...· 0 citations
We introduce the orbital cluster expansion (OCE), a linear regression on physics-motivated local features derived from atomic orbital eigenenergies, and benchmark it against the SPICE 2.0 biomolecular data set at the ωB97M-D3BJ/def2-TZVPPD level. With regression of formation energies on 677 dipeptides spanning the natural amino acids, ridge regression on 414 OCE features attains a parent-stratified test root-mean-square error of 30 meV per atom with Spearman ρ = 0.97 and R 2 = 0.95 against a target spread of only 0.13 eV per atom, matching MACE-OFF23(L) and ANI-2x trained with 104–106 conformations but with ∼103 fewer training points. Comparable accuracy holds on 500 PubChem drug-like molecules and 500 DES370K dimers. We characterize a fundamental dual regime: intermolecular ranking is preserved across chemistries, while intraconformer ranking is random because the basis cannot resolve geometry-only variation within a fixed connectivity. OCE is a transparent, physically interpretable surrogate for intermolecular biomolecular screening.
D. L. Azevedo· Journal of Physical Chemistr...· 0 citations
This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.
Jigyasa Nigam, T. Smidt, G. Dusson· Journal of Chemical Physics· 2 citations
Accurate prediction of electronic Hamiltonians would enable broad property inference while avoiding the high computational cost of Density Functional Theory (DFT). However, progress toward general-purpose materials foundation models is limited by a data bottleneck: existing Hamiltonian datasets are typically small, lack structural diversity, and often omit essential relativistic physics such as spin--orbit coupling (SOC). We therefore construct UniHam, a large-scale Hamiltonian dataset and benchmark suite comprising 100,000+ DFT-computed complex-valued Hermitian Hamiltonians with full SOC, covering 72 elements and a wide range of crystal geometries and symmetries (spanning diverse lattice types and space-group families). Building on UniHam, we benchmark two representative state-of-the-art models under a standardized protocol and introduce complementary evaluation metrics that jointly assess three dimensions: (i) Hamiltonian reconstruction accuracy, (ii) out-of-distribution (OOD) generalization across composition/symmetry shifts, and (iii) the ability to support downstream property prediction from the predicted Hamiltonians. Experiments on UniHam demonstrate that the proposed benchmark and metrics effectively differentiate model capabilities, revealing intrinsic SOC- and element-dependent failure modes, large variations in compositional OOD robustness, and the necessity of spectral-level evaluation to assess whether Hamiltonian predictions reliably support downstream electronic-structure properties. Overall, UniHam provides a reproducible, SOC-complete benchmark that can sharpen model comparisons and accelerate the development of next-generation foundation models for quantum materials.
Yuewen Huang, Pin Chen, Yutong Lu· Proceedings of the 32nd ACM...· 0 citations
Results indicate that the topology-aligned inductive bias is the active ingredient driving parameter efficiency at QM9 scale, with implications for matched-baseline benchmarking in quantum machine learning.