Skip to content

Molecular Property Prediction via Sparse Binary Matrix Representation and Convolutional Neural Networks

Aug 2026 · Industrial & Engineering Chemistry Research · 0 citations · 78 references

TL;DR

The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model.

Abstract

A simple and interpretable matrix-based representation is presented for predicting molecular properties, specifically individual HOMO and LUMO frontier orbital energies and their resulting energy gaps, of functionalized organic molecules using a Convolutional Neural Network (CNN). Each molecule is encoded as a sparse binary matrix (SBMR) that captures the identity and position of substituents on a fixed molecular backbone. The model was initially benchmarked across four molecular families: n-butane, i-butane, cyclobutadiene, and quinone, achieving a combined RMSE of 4.0 kcal mol–1 for gap predictions compared to DFT-computed references, with over 85% of predictions falling within ±5% error. To contextualize this performance, the model was benchmarked against six established featurization methods spanning 2D topology and 3D physics-based approaches: the Coulomb Matrix (CM), Smooth Overlap of Atomic Positions (SOAP), 2D and 3D Message-Passing Neural Networks (MPNN), Random Forest with Morgan Fingerprints (RF-MF), and Uni-Mol+. The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model. This is achieved while offering distinct advantages through a dramatically smaller feature space and less stringent input data requirements. To demonstrate extensibility to complex catalytic systems, the architecture was applied to a combinatorial data set of 1,4-dihydropyridine derivatives, a class of redox mediators utilized in electrochemical and biochemical applications. For these highly functionalized heterocycles, the model successfully decoupled the energy gap into its constituent levels, predicting HOMO and LUMO energies with an RMSE of 2.8 and 2.5 kcal mol–1, respectively. The resulting framework couples high predictive accuracy with representational interpretability, offering a transparent and customizable tool for property prediction with direct applications in molecular screening, rational design, and electrocatalyst optimization.

View source

Similar papers

Open access Jul 2026

Sparse Linear Surrogates Match Neural Network Potentials on the SPICE Biomolecular Benchmark with Three Orders of Magnitude Smaller Training Sets

We introduce the orbital cluster expansion (OCE), a linear regression on physics-motivated local features derived from atomic orbital eigenenergies, and benchmark it against the SPICE 2.0 biomolecular data set at the ωB97M-D3BJ/def2-TZVPPD level. With regression of formation energies on 677 dipeptides spanning the natural amino acids, ridge regression on 414 OCE features attains a parent-stratified test root-mean-square error of 30 meV per atom with Spearman ρ = 0.97 and R 2 = 0.95 against a target spread of only 0.13 eV per atom, matching MACE-OFF23(L) and ANI-2x trained with 104–106 conformations but with ∼103 fewer training points. Comparable accuracy holds on 500 PubChem drug-like molecules and 500 DES370K dimers. We characterize a fundamental dual regime: intermolecular ranking is preserved across chemistries, while intraconformer ranking is random because the basis cannot resolve geometry-only variation within a fixed connectivity. OCE is a transparent, physically interpretable surrogate for intermolecular biomolecular screening.

D. L. Azevedo · 0 citations
Preprint Jul 2026

Graph Neural Network Force Fields (GPTFF-mol) for Organic Molecules from Optimization Trajectories (OpenGEM26)

Density functional theory (DFT) serves as a reliable tool for atomistic molecular simulations, while machine learning potentials have become powerful complements to balance accuracy and efficiency. In this work, we release OpenGEM26 (Open Generated Ensemble of Molecules, 2026), a large-scale dataset comprising 200,000 unique molecules and 4.4 million conformations composed of H, C, N, O, S and Cl with up to ten heavy atoms. All calculations are carried out at the {\omega}B97X-D/Def2-SVP and Def2-TZVP levels with dispersion corrections, and complete structural optimization trajectories and abundant non-equilibrium structures are recorded. Statistical analyses confirm that this dataset covers a broader conformational space than QM9 in terms of energy, bond lengths and bond angles. A graph neural network-based potential GPTFF-mol is trained using the new dataset, achieving an energy mean absolute error of 16 meV/molecule, which is equivalent to 0.82meV/atom, and superior force prediction performance compared with ANI-2x. Validated by butane rotation and keto-enol tautomerization tests, the model accurately describes molecular dynamical behaviors and reaction barriers at distorted geometries. This work provides a high-quality resource and robust ML potential for efficient simulations of sulfur- and chlorine-containing organic molecules.

Yifan Huang, Fankai Xie, Jiangnan Zheng et al. · 0 citations
Open access Jul 2026

Path-weighted atom vectors and ChemBERTa fusion for predicting physicochemical properties

Molecular property prediction is central to cheminformatics and environmental chemistry, where accurate modeling of physicochemical properties supports risk assessment and molecular design. Classical descriptors and recent advances such as ChemBERTa have enabled learning chemically contextual representations directly from SMILES, while the integration of structured descriptors with transformer-based embeddings offers a promising pathway toward accurate and interpretable prediction. In this study, we introduce Path-Weighted Atom Vectors (PWAVs), a descriptor family that captures atom-level, environment-aware structural information. We evaluate PWAV both as a standalone representation and in combination with ChemBERTa embeddings through a gated fusion architecture incorporating modality dropout, FiLM conditioning, and auxiliary supervision. Experiments on six physicochemical property datasets ( log P, log S, log BCF, boiling point, melting point, and vapor pressure) show that PWAV generally improves over classical fingerprint descriptors within learned models and achieves competitive performance relative to established external baselines on several endpoints. The strongest gains are observed for boiling point, aqueous solubility, and partition coefficient prediction, where descriptor-embedding fusion yields the best results among the learned models considered. Ablation analyses demonstrate that PWAV contributes complementary structural information beyond SMILES-only ChemBERTa representations, while SHapley Additive exPlanations-based interpretability shows that predictive signal is concentrated within a compact subset of features, enabling an efficient reduced representation (PWAV-64). Nested cross-validation further confirms the robustness of PWAV within the XGBoost framework. Overall, PWAV provides a compact, interpretable, and extensible descriptor framework that integrates effectively with modern representation-learning approaches. These results position PWAV as a competitive and chemically transparent component for hybrid molecular property prediction, rather than as a replacement for domain-specific benchmark systems.

M. Afzal, S. Siddiqi · 0 citations
Jul 2026

tmGNN-XAI: An Explainable Graph Neural Network Tool for Predicting Electronic Properties of Transition Metal Complexes from SMILES.

Predicting the electronic properties of transition metal complexes (TMCs) from 2D molecular graphs remains challenging; organic-trained property models lack TMC transferability, universal interatomic potentials require 3D coordinates rather than SMILES, and tools providing holistic electronic property prediction with atom-level explainability and calibrated uncertainty remain limited. We present tmGNN-XAI, a multitask relational graph convolutional network that predicts seven quantum-chemical properties of TMCs directly from SMILES strings and produces perturbation-based atom-level attributions for each prediction. The model encodes dative coordination bonds as a dedicated edge type distinct from covalent bonds and is trained on 100,703 complexes from the tmQM data set spanning 30 transition metals. Test-set performance is competitive with a Chemprop D-MPNN baseline, achieving R2 = 0.979 for metal partial charge and R2 = 0.964 and 0.949 for HOMO and LUMO energies. Across all 100,703 complexes, donor atoms (N, O, S, P) appear among the top-five most important atoms in more than 99.8% of complexes for every property, a large-scale data-driven result consistent with ligand field theory. A trust framework combining ensemble agreement with attribution direction separates predictions into four reliability scenarios; confident predictions achieve 1.6 to 2.5 times lower mean absolute error than uncertain ones for five of seven properties. The framework generalizes to cross-level DFT validation, phototherapy candidate screening (area under the ROC curve (AUC) = 0.735), and indirect redox prediction via Koopmans' theorem. An interactive web application makes property predictions, atom-level attributions, and trust labels accessible without programming or DFT expertise. tmGNN-XAI is designed as an explainable, first-tier screening tool for TMC electronic property estimation.

Abdulmujeeb T. Onawole · 2 citations
Preprint Jul 2026

Multimodal Molecular Representation Learning with Graph Neural Networks, Deep&Cross Networks, and SMILES Embeddings

Molecular property prediction often relies on isolated data modalities, where continuous 3D graph neural networks (GNNs) struggle to efficiently capture long-range topological dependencies and exact macroscopic heuristics. In this work, we introduce a parameter-efficient Tri-Branch Modular Fusion Neural Network that synthesizes three orthogonal modalities: 3D spatial geometry (SchNet), discrete topological grammar (SMILES via ChemBERTa), and explicit macroscopic physicochemical descriptors (Deep&Cross Network). By bypassing standard scalar readouts and employing a shared late-fusion architecture, the framework establishes a mathematically rigorous multimodal latent space that effectively resolves the arithmetic and oversmoothing limitations of local message passing. We evaluate the proposed architecture on the QM9 benchmark, targeting the extensive thermodynamic property of atomization energy at 0 K ($U_0^{\mathrm{atom}}$). Through systematic combinatorial ablation and latent bottleneck optimization ($d_e=64$), the tri-modal framework achieves a validation Mean Absolute Error (MAE) of 0.0207 eV. Operating with fewer than one million parameters, this architecture decisively surpasses the sub-chemical accuracy threshold and yields a substantial 20.6% error reduction over a strictly controlled geometric baseline. Ultimately, our findings demonstrate that integrating orthogonal macroscopic and topological data streams provides a synergistic, $\mathcal{O}(1)$ physical shortcut. This multimodal alignment offers a highly efficient alternative to brute-force parameter scaling, establishing a robust surrogate model for high-throughput virtual screening (HTVS) pipelines.

Qiwei Han, Chi Zhou, Ruobing Wang et al. · 0 citations
Preprint Jul 2026

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordinate reconstruction, and pairwise distance reconstruction. On the CREMP-CycPeptMPDB dataset, training EnsembleEGNN from scratch fails entirely ($R^2=0.005$). However, the pretrained model reaches $R^2=0.477$ and Pearson $r=0.699$, outperforming the sequence-only BERT baseline ($R^2=0.439$, Pearson $r=0.667$). When EnsembleEGNN is co-trained end-to-end with the BERT sequence encoder, the hybrid model improves further to $R^2=0.538$ and Pearson $r=0.737$. These results demonstrate that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.

Aaron L. Feller, Kristine Deibler, Maxim Secor · 0 citations