Skip to content
Open access

MolDeTr: A Chemistry-Informed Deep Learning Model for Next-Generation Automated Analysis of 1H NMR Spectra

Aug 2026 · Analytical Chemistry · Vol 98, pp. 23399 - 23412 · 1 citation · 62 references
Medicine

TL;DR

MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches.

Abstract

Accurate interpretation of one-dimensional proton nuclear magnetic resonance (1H NMR) spectra remains a rate-limiting step in molecular structure elucidation, particularly when signal overlap, strong spin coupling, and instrumental distortions mask key features. Existing automated approaches depend on computationally intensive and sensitive iterative quantum-mechanical fitting and still require expert oversight. Here we introduce MolDeTr, a chemistry-informed deep-learning framework derived from the detection-transformer architecture that unifies peak picking, multiplet identification, and extraction of chemical shifts, scalar coupling constants, relaxation-dependent decay times, and proton counts in a single-network pass. The method targets prototypical spin systems of single-component small molecules in 1D 1H NMR, with up to ten distinct groups of chemically equivalent spins (multiplets). MolDeTr is trained exclusively on synthetic spectra generated by spin-dynamics simulations and augmented with realistic experimental artifacts, enabling it to generalize to unseen compoundsincluding experimental spectra with overlapping and strongly coupled multipletswithout reference standards or prior spin-system knowledge. Unlike structure-conditioned shift-prediction or calculation models, e.g., density functional theory (DFT), that assume the molecular structure is known, MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches. Benchmarking against a diverse experimental set of 1H NMR spectra of modestly sized small molecules, spanning 80 to 600 MHz base frequency, shows median absolute errors of 0.89 Hz for chemical shifts and 0.20 Hz for coupling constants, while absolute proton counts are predicted with 93.5% accuracy, outperforming state-of-the-art spectrum analysis software and experienced spectroscopists. By eliminating iterative fitting and expert intervention, MolDeTr offers a scalable route to fully automated spectral analysis, accelerating molecular discovery across the chemical sciences.

Read PDF

Similar papers

Aug 2026

PINS: A Physics-Informed Generative Framework for De Novo Structure Elucidation from 1D NMR Spectra.

PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors, provides a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.

Pengfei Liu, Cuimei Liu, Laiqun Xia et al. · 0 citations
Preprint Jul 2026

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.

Chengchun Liu, Zhiyuan Yan, Li Yuan et al. · 0 citations
Preprint Jul 2026

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem

The results show that reframing NMR elucidation as an LLM-guided constrained search, rather than a modeling task, yields substantial gains and suggests a path toward multi-step orchestration frameworks that integrate a variety of tools, models, and domain knowledge to assist in automating spectroscopic analysis.

I. Morales, Damon J. Hinz, Marvin Alberts et al. · 0 citations
Aug 2026

Improving Quantum-Chemical Prediction of 19F NMR Chemical Shifts via Machine Learning

19F nuclear magnetic resonance (NMR) spectroscopy is widely used for structural elucidation of fluorinated molecules, but reliable prediction of 19F chemical shifts remains challenging because quantum-chemical calculations are computationally demanding and their accuracy can vary across diverse molecular environments. In this work, we develop a machine-learning-assisted framework to improve quantum-chemical prediction of 19F NMR chemical shifts. A data set of 2605 experimental shifts was compiled from the literature, and isotropic shielding constants were calculated using a density functional theory (DFT)/gauge-including atomic orbital (GIAO) calculation protocol. Machine learning was then used to analyze the relationship between calculated shielding values and experimental chemical shifts. The analysis indicates that the data set can be partitioned into operationally defined, structure-associated regimes in which the mapping between calculated shielding and experimental shift differs systematically. By identifying the structural characteristics of these regimes and constructing prediction models separately for each region, the overall predictive accuracy of the quantum-chemical framework is significantly improved. The resulting models achieve mean absolute errors below 4 ppm and show practical promise under the tested benchtop 60 MHz conditions after simple linear calibration. Application to a fluorinated reaction mixture further demonstrates the utility of the approach for assisting spectral interpretation and prioritizing candidate structures. These results show that the main contribution of the present work is not simply applying machine learning to 19F NMR prediction, but using machine learning to diagnose and correct subset-dependent limitations in the shielding-shift relationship within a practical quantum-chemical workflow.

Dongdong Chen, Yuanxiang Ye, Yijie Zhu et al. · 0 citations
Open access Aug 2026

Resolving Chemically Inequivalent 11B NMR Sites via Interpretable Hybrid Machine Learning

A manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally.

Penghui Li, Ben Gao, Shiyang Wang et al. · 0 citations
Preprint Aug 2026

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

MACROS establishes a scalable foundation for fully automated structure elucidation, and catalyzes accelerated molecular discovery toward autonomous laboratories, and augments chemists via collaboration to deliver sixfold faster, 40% more accurate elucidation.

Bingsen Xue, Zhuojun Jiang, Jianhao Zhang et al. · 0 citations