Aug 2026· Nature Computational Science· 0 citations· 49 references
Medicine
TL;DR
Self-supervised pretraining substantially improves TS prediction for previously unseen systems, lowering the median root-mean-square deviation of TS geometries on Transition1x-TMC reactions and reducing fine-tuning data requirements, enabling reliable performance even in low-data regimes.
This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 0 citations
A framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space is introduced, establishing a blueprint for robust machine learning in synthetic chemistry.
Paulo Neves, Bo Hao, Santeri Aikonen et al.· Nature Computational Science· 1 citation
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
Idil Ismail, Gregory A Landrum, Sereina Riniker· Journal of the American Chem...· 0 citations
ChemFusion is presented, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates and reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.
A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.
Xuefeng Li, Haoke Qiu, Hanwen Pei et al.· Journal of Physical Chemistr...· 0 citations
This work not only establishes a pioneering paradigm for interpretable ML-driven force field refinement but also provides the first feature engineering solution incorporating chemical, physical, and structural information specifically designed for the machine learning of energetic molecular crystals.
Qi He, Pengju Wang, Xudong He et al.· Molecules· 0 citations