Skip to content

Charting the small-molecule universe from mass spectra with neuro-symbolic AI

Aug 2026 · bioRxiv · 0 citations · 72 references
Biology

TL;DR

AIMe (AI Molecule Explorer), a multi-agent neuro-symbolic AI framework that transforms the interpretation of unknown spectra into an omics-scale exploration across the known structural space, providing chemically interpretable annotations, is introduced.

View source

Similar papers

Preprint Jul 2026

MARLIN: De Novo Molecular Structure Elucidation from Tandem Mass Spectra without a Ground-Truth Formula

Untargeted tandem mass spectrometry (MS/MS) detects thousands of small molecules per biological sample, yet most go unidentified because they are absent from spectral libraries. These uncharacterized metabolites and natural products are precisely the compounds that matter for drug discovery, biomarker research, and exposomics. Computational de novo structure elucidation could close this gap, but almost all state-of-the-art methods assume the ground-truth molecular formula is known, an oracle that does not exist for genuinely novel compounds and is itself predicted with substantial error. We present MARLIN, a de novo method that elucidates structures directly from a spectrum with no molecular formula at any stage. A self-supervised encoder predicts a molecular fingerprint from the raw peaks, and a block-diffusion language model generates candidate structures conditioned only on the fingerprint and the instrument-measured precursor mass. A provably safe mass-shell constraint keeps every candidate consistent with the measured mass without fixing the atom inventory, and candidates are accepted by exact parts-per-million mass agreement. A symmetric noise objective absorbs encoder error, and a candidate-diversity mechanism keeps the candidates from collapsing to a single structure. On the NPLIB1 benchmark, MARLIN is the strongest method evaluated without a ground-truth formula across exact-match accuracy, structural distance, and fingerprint similarity, and it recovers the correct molecular formula as a byproduct about as often as a dedicated predictor without ever using one. MARLIN enables reliable de novo structure elucidation in the realistic discovery regime where the molecular formula is unavailable.

Xujun Che, Xiuxia Du, Depeng Xu · 0 citations
Jul 2026

Can Large Language Models Translate MS/MS into Molecular Caption?

Mass spectrometry (MS) is central to molecular discovery, yet the interpretation of tandem mass spectra (MS/MS) remains limited by database dependence and an incomplete structural resolution. Here, we explore whether large language models (LLMs) can directly translate MS/MS spectra to chemically meaningful molecular descriptions. We introduce MS2LLM, a framework that represents spectra and molecular structures as natural languages and learns their correspondence through instruction tuning. MS2LLM generates hierarchical molecular descriptions, including functional groups, substructures, and chemical classes, directly from the spectral input. Across multiple datasets, it outperforms general-purpose LLMs and conventional spectral learning methods in both descriptive accuracy and chemical classification. Importantly, the model produces interpretable outputs that capture structural semantics rather than exact structures, offering a complementary paradigm for structural inference of unknowns.

Menglin Zhou, Chenhao Gong, Shan Cong et al. · 0 citations
Aug 2026

Prediction of mass spectra using large chemical language models and verification of adaptability in data-scarce domains

Machine learning methods for predicting the electron ionization mass spectra from molecular structures have shown promise for environmental chemical identification, but their performance under domain-specific data scarcity remains poorly understood. We systematically compare a conventional multilayer perceptron model (NEIMS) with a Transformer-based chemical foundation model (MolFormer-XL) for the electron ionization mass spectrometry spectrum prediction under controlled few-shot conditions. Using fluorine-containing molecules as a broader proxy domain, including a PFAS-like subset, motivated by the practical challenge of detecting novel fluorinated contaminants with limited reference data, we vary the number of domain-specific training examples from 5 to 175 while maintaining fixed validation and test sets. Across all few-shot conditions and three of four evaluation metrics (weighted cosine similarity, intensity-weighted precision, and top-10 precision), MolFormer-XL consistently outperforms NEIMS, while intensity-weighted recall remains comparable between the two models. The largest performance gaps are observed in extreme data-scarcity regimes. These results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.

Satoki Muto, Akiko Kumada, Masahiro Sato · 0 citations
Preprint Aug 2026

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

MACROS establishes a scalable foundation for fully automated structure elucidation, and catalyzes accelerated molecular discovery toward autonomous laboratories, and augments chemists via collaboration to deliver sixfold faster, 40% more accurate elucidation.

Bingsen Xue, Zhuojun Jiang, Jianhao Zhang et al. · 0 citations
Open access Aug 2026

Learning from human and chemical languages to predict biological function

PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure, establishes that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.

Clayton W. Kosonocky, Nikol Kadeřábková, Kangsan Kim et al. · 0 citations
Open access Aug 2026

From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation

Untargeted metabolomics often results in a significant portion of unannotated metabolites, or “metabolic dark matter,” which hinders biological interpretation. A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC–MS/MS data from pregnant women with obesity as a biologically relevant test dataset. The first step involved clustering 1,021 known metabolites into ten structurally coherent groups based on the Tanimoto similarity, thus defining the biologically relevant chemical space of the dataset. These metabolites were further characterized by Absorption, Distribution, Metabolism, and Excretion (ADME) profiling, protein target prediction, molecular docking and Kyoto Encyclopedia of Genes and Genomes pathway mapping analysis, to establish biological plausibility and functional perspective. Candidate structures for 1,836 unannotated features were retrieved from PubChem using molecular formula and molecular weight matching within a ±0.5 Da tolerance. This search yielded 569,115 candidate structures, of which 368,197 unique structures were retained after curation. Tanimoto coefficient filtering reduced the candidate pool to 19,868 structurally plausible candidates, and retention time-based prioritization further refined this set to 418 high confidence candidate annotations, including 83 database-supported candidates identified through HMDB and LIPID MAPS structure database cross-referencing. RT-based prioritization effectively distinguished positional isomers sharing the same molecular formula by incorporating agreement between predicted and experimentally observed retention times. This improved discrimination among structurally similar candidates, expanded metabolite annotation confidence, and provided a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics. Clustered-based workflow integrating chemical similarity and retention time to prioritize and annotate unknown metabolites

D. Bhandari, H. Paz, Keith Henderson et al. · 0 citations