Aug 2026· Journal of Chemical Theory and Computation· 1 citation· 79 references
Abstract
Covalent organic frameworks (COFs) are highly ordered, porous organic materials whose reticular construction from tailored nodes and linkers enables atomic-level control over structure and function. The design space of COFs is vast with virtually unlimited combinations of nodes, linkers, and functional groups. Interpretable machine learning (ML) offers a pathway to navigate this complexity by identifying the structural features that govern materials performance, yet interpretability often comes at the cost of predictive accuracy. In this work, we introduce a novel multiple-kernel learning framework that achieves both accuracy and mechanistic insight. A multiple-kernel ridge regression (MKRR) model was trained on band gaps predicted from GFN1-xTB level theory for a data set of 232 theoretical pyranoazacoronene (PAC) COFs produced from eight different conjugated linkers and 29 functional groups. Modifying these building units alone produced a range of band gaps between 0.4–2 eV. Manual analysis of the theoretical band gaps versus the linker indicates that breaking the conjugation pathway by altering the bond angle or by introducing a σ-bond increases the band gap while increasing the length of the linker decreases the band gap. All functional groups appear to reduce the band gap with three specific electron withdrawing groups reducing the band gap near 0.4 eV. For the ML, the building units were represented with three independent kernels that encoded the local environments of each node, linker, and functional group calculated from the Smooth Overlap of Atomic Positions (SOAP). After decomposing each kernel’s contribution to the model’s global predictions, we found that the MKRR model successfully captures the underlying structure–property relationships that influence the band gap. These results demonstrate that MKRR is an effective and interpretable framework for understanding and designing functional COFs.
A semiempirical extended tight-binding approach (GFN1-xTB) is employed to compute the electronic properties of a dataset of MOFs, and it is shown that GFN1-xTB approximates MOF band gaps well, as compared to semilocal DFT.
A. Jose, A. Walsh· Journal of Chemical Theory a...· 0 citations
A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.
Xuefeng Li, Haoke Qiu, Hanwen Pei et al.· Journal of Physical Chemistr...· 0 citations
The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model.
Abdulaziz W. Alherz, C. Tezak, Mohammed S. Alhajeri· Industrial & Engineering...· 0 citations
Metal-organic frameworks (MOFs) and MOF-like porous materials exhibit vast structural diversity and support critical applications in gas storage, separations, and catalysis. Predictive modeling remains difficult because their structure-property relationships are multiscale and cage-like, governed by both local chemical environments and global pore-network topology. These challenges, together with sparse and unevenly distributed labeled data, hinder generalization across material families. We develop an interaction topology theory and propose the interaction topological transformer (ITT), a data-efficient framework that captures materials information across multiple scales and levels, including structural, elemental, atomic, and pairwise-elemental organization. ITT extracts scale-aware features reflecting both compositional and relational structures in complex porous frameworks and integrates them through a transformer architecture for joint reasoning across scales. Using self-supervised pretraining on more than 0.6 million unlabeled structures followed by supervised fine-tuning, ITT achieves accurate, transferable, state-of-the-art predictions for adsorption, transport, and stability properties across 17 tasks, providing a principled and scalable strategy for learning-guided discovery in diverse MOF-like materials.
We introduce the orbital cluster expansion (OCE), a linear regression on physics-motivated local features derived from atomic orbital eigenenergies, and benchmark it against the SPICE 2.0 biomolecular data set at the ωB97M-D3BJ/def2-TZVPPD level. With regression of formation energies on 677 dipeptides spanning the natural amino acids, ridge regression on 414 OCE features attains a parent-stratified test root-mean-square error of 30 meV per atom with Spearman ρ = 0.97 and R 2 = 0.95 against a target spread of only 0.13 eV per atom, matching MACE-OFF23(L) and ANI-2x trained with 104–106 conformations but with ∼103 fewer training points. Comparable accuracy holds on 500 PubChem drug-like molecules and 500 DES370K dimers. We characterize a fundamental dual regime: intermolecular ranking is preserved across chemistries, while intraconformer ranking is random because the basis cannot resolve geometry-only variation within a fixed connectivity. OCE is a transparent, physically interpretable surrogate for intermolecular biomolecular screening.
D. L. Azevedo· Journal of Physical Chemistr...· 0 citations
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.