Aug 2026· Nature Computational Science· 1 citation· 29 references
Medicine
TL;DR
A framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space is introduced, establishing a blueprint for robust machine learning in synthetic chemistry.
Abstract
The Buchwald-Hartwig cross-coupling is a cornerstone of modern pharmaceutical synthesis, yet predictive modeling of its outcomes remains constrained by data quality and chemical space coverage. Electronic laboratory notebooks contain heterogeneous, noisy records, while open-source high-throughput experimentation (HTE) datasets are fragmented and narrow in scope, leading to poor model performance on unseen substrates and conditions. Here we introduce a framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space. By merging published Buchwald-Hartwig HTE data with new experimental results, we achieve a model with predictive power across novel substrates and conditions, delivering improved out-of-distribution predictions compared with previous approaches. Crucially, model-guided reagent recommendations were validated experimentally, confirming the framework's utility to uncover unexplored reactivity. This work establishes a blueprint for robust machine learning in synthetic chemistry and enables preemptive in silico reagent screening to accelerate pharmaceutical discovery.
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
Idil Ismail, Gregory A Landrum, Sereina Riniker· Journal of the American Chem...· 0 citations
ChemFusion is presented, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates and reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.
Retrosynthesis, the process of predicting reactants from products, remains a critical challenge in computational chemistry and drug discovery. While recent deep learning methods have shown strong performance, they remain overly reliant on reaction datasets, which are limited in availability and quality. Large-scale unlabeled molecular data encode rich structural patterns that can be leveraged to learn transferable chemical knowledge, but remain largely unexplored. In this work, we propose KnowRetro (Knowledge-Guided Retrosynthesis Prediction), a chemically-aware framework that learns chemical knowledge from large-scale unlabeled molecules to enhance the accuracy and diversity of retrosynthesis prediction. Specifically, KnowRetro first builds a hierarchical knowledge graph from millions of unlabeled molecules, which captures transformation-relevant relationships among molecules, substructures, and functional groups. It then employs chemically guided pre-training based on substructure decomposition to encourage the model to capture fundamental reaction patterns, followed by fine-tuning with an adapter designed to inject task-relevant knowledge into reactant generation. Extensive experiments demonstrate that KnowRetro achieves high accuracy with improved robustness and diversity in reactant generation. Our code is available at https://github.com/chenyujie1127/KnowRetro.
Yujie Chen, Tengfei Ma, Zhou Yu et al.· Proceedings of the 32nd ACM...· 0 citations
Novo-1, a coarse-grained cofolding framework for binding- affinity prediction, offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, and demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets.
Nikhil Shenoy, David Errington, Emmanuel Bengio et al.· bioRxiv· 0 citations
This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 0 citations
This work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles, and introduces a suite of ADMET models for kinetic solubility, lipophilicity, and Caco-2 permeability to improve developability at the point of selection.