Aug 2026· Analytical Chemistry· Vol 98, pp. 23399 - 23412· 1 citation· 62 references
Medicine
TL;DR
MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches.
Abstract
Accurate interpretation of one-dimensional proton nuclear magnetic resonance (1H NMR) spectra remains a rate-limiting step in molecular structure elucidation, particularly when signal overlap, strong spin coupling, and instrumental distortions mask key features. Existing automated approaches depend on computationally intensive and sensitive iterative quantum-mechanical fitting and still require expert oversight. Here we introduce MolDeTr, a chemistry-informed deep-learning framework derived from the detection-transformer architecture that unifies peak picking, multiplet identification, and extraction of chemical shifts, scalar coupling constants, relaxation-dependent decay times, and proton counts in a single-network pass. The method targets prototypical spin systems of single-component small molecules in 1D 1H NMR, with up to ten distinct groups of chemically equivalent spins (multiplets). MolDeTr is trained exclusively on synthetic spectra generated by spin-dynamics simulations and augmented with realistic experimental artifacts, enabling it to generalize to unseen compoundsincluding experimental spectra with overlapping and strongly coupled multipletswithout reference standards or prior spin-system knowledge. Unlike structure-conditioned shift-prediction or calculation models, e.g., density functional theory (DFT), that assume the molecular structure is known, MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches. Benchmarking against a diverse experimental set of 1H NMR spectra of modestly sized small molecules, spanning 80 to 600 MHz base frequency, shows median absolute errors of 0.89 Hz for chemical shifts and 0.20 Hz for coupling constants, while absolute proton counts are predicted with 93.5% accuracy, outperforming state-of-the-art spectrum analysis software and experienced spectroscopists. By eliminating iterative fitting and expert intervention, MolDeTr offers a scalable route to fully automated spectral analysis, accelerating molecular discovery across the chemical sciences.
PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors, provides a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
Chengchun Liu, Zhiyuan Yan, Li Yuan et al.· 0 citations
The results show that reframing NMR elucidation as an LLM-guided constrained search, rather than a modeling task, yields substantial gains and suggests a path toward multi-step orchestration frameworks that integrate a variety of tools, models, and domain knowledge to assist in automating spectroscopic analysis.
I. Morales, Damon J. Hinz, Marvin Alberts et al.· 0 citations
A machine-learning-assisted framework to improve quantum-chemical prediction of 19F NMR chemical shifts by using machine learning to diagnose and correct subset-dependent limitations in the shielding-shift relationship within a practical quantum-chemical workflow is developed.
Dongdong Chen, Yuan-Xiang Ye, Yijie Zhu et al.· Journal of Chemical Informat...· 0 citations
A manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally.
Penghui Li, Ben Gao, Shiyang Wang et al.· JACS Au· 0 citations
MACROS establishes a scalable foundation for fully automated structure elucidation, and catalyzes accelerated molecular discovery toward autonomous laboratories, and augments chemists via collaboration to deliver sixfold faster, 40% more accurate elucidation.
Bingsen Xue, Zhuojun Jiang, Jianhao Zhang et al.· 0 citations