A manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally.
Abstract
Organoboron compounds are widely used across pharmaceuticals and materials science, where 11B NMR spectroscopy serves as a valuable tool for structural characterization. However, severe spectral line broadening induced by the quadrupolar nature of the boron nucleus often causes signal overlap, making it exceptionally difficult to experimentally resolve chemically inequivalent sites in complex multiboron architectures. While traditional density functional theory can resolve these ambiguities, it faces prohibitive computational bottlenecks, whereas data-driven alternatives remain constrained by the scarcity of high-quality data sets. Herein, we report a manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow. Interpretable machine learning identifies a strong correlation between the BCUT2D_MRLOW descriptor and the boron hybridization. Integrating these ML-derived features as prior knowledge, we developed a prior-guided Graph Transformer for accurate atom-level chemical shift prediction. Notably, the model provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally. We further deploy the framework as an open-access Web tool to support the rapid structural analysis of organoboron compounds.
19F nuclear magnetic resonance (NMR) spectroscopy is widely used for structural elucidation of fluorinated molecules, but reliable prediction of 19F chemical shifts remains challenging because quantum-chemical calculations are computationally demanding and their accuracy can vary across diverse molecular environments. In this work, we develop a machine-learning-assisted framework to improve quantum-chemical prediction of 19F NMR chemical shifts. A data set of 2605 experimental shifts was compiled from the literature, and isotropic shielding constants were calculated using a density functional theory (DFT)/gauge-including atomic orbital (GIAO) calculation protocol. Machine learning was then used to analyze the relationship between calculated shielding values and experimental chemical shifts. The analysis indicates that the data set can be partitioned into operationally defined, structure-associated regimes in which the mapping between calculated shielding and experimental shift differs systematically. By identifying the structural characteristics of these regimes and constructing prediction models separately for each region, the overall predictive accuracy of the quantum-chemical framework is significantly improved. The resulting models achieve mean absolute errors below 4 ppm and show practical promise under the tested benchtop 60 MHz conditions after simple linear calibration. Application to a fluorinated reaction mixture further demonstrates the utility of the approach for assisting spectral interpretation and prioritizing candidate structures. These results show that the main contribution of the present work is not simply applying machine learning to 19F NMR prediction, but using machine learning to diagnose and correct subset-dependent limitations in the shielding-shift relationship within a practical quantum-chemical workflow.
Dongdong Chen, Yuanxiang Ye, Yijie Zhu et al.· Journal of Chemical Informat...· 0 citations
MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches.
N. Schmid, Marc Wanner, G. Fischetti et al.· Analytical Chemistry· 1 citation
Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs. This limits model predictions to isomer identification rather than full molecular structure prediction. Although transformer models have been shown to identify molecular isomers with high accuracy, their reliability for unconstrained structure elucidation is comparatively low and poorly understood. In this work, we propose and evaluate key modifications to the traditional encoder-decoder transformer. To better address the vast chemical space of the unconstrained problem, we implement a novel Mixture-of-Experts (MoE) decoder module that utilizes non-additive aggregation via linear-order statistics and the Choquet integral. We further modify the transformer to utilize these non-additive operators when aggregating spectral representations as well. Together with an auxiliary contrastive alignment loss term, these enhancements improve Top-K prediction accuracy by over 10 percentage points compared to baseline IR-only models. Through sub-structure fragment analysis of molecular predictions, we further confirm that infrared spectra encode the vast majority of relevant chemical information, implying that the higher performance of isomer-ranking models is largely due to underrepresented or overlapping absorption bands for molecules in the explored chemical space. Ultimately, by demonstrating the efficacy of automated molecular structure elucidation from measured IR spectra, this work serves to significantly broaden the utility of AI in analytical chemistry.
Ethan J. Mick, C. Sweet, M. Young et al.· 0 citations
The results show that reframing NMR elucidation as an LLM-guided constrained search, rather than a modeling task, yields substantial gains and suggests a path toward multi-step orchestration frameworks that integrate a variety of tools, models, and domain knowledge to assist in automating spectroscopic analysis.
I. Morales, Damon J. Hinz, Marvin Alberts et al.· 0 citations
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predictors use undirected molecular graphs and pairwise edges. This choice misses asymmetric interactions, nonreversible dynamics, and motif level effects from functional groups and ring systems. We propose ChemHyperMag for multitask ADMET prediction under missing labels. ChemHyperMag builds a functional group hypergraph from rings, BRICS fragments, Bemis-Murcko scaffolds, and bonds. It also defines a potential driven nonreversible flow guided by electronegativity and Gasteiger partial charges. The resulting circulation is encoded by a Hermitian magnetic Laplacian and processed with a magnetic Chebyshev encoder. We perturb magnetic phases to form stochastic views and train with an InfoNCE objective. Experiments on multiple ADMET benchmarks show improvements over recent methods with fewer labeled samples and no conformers. ChemHyperMag is scalable and provides interpretable directional signals through its magnetic phases.
Hexiao Ding, Hongzhao Chen, Jing Lan et al.· 0 citations
PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors, provides a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.