Skip to content
Open access

Evaluation of Machine Learning Models for Condition Optimization in Diverse Amide Coupling Reactions

Jul 2026 · ACS Omega · Vol 11, pp. 41886 - 41895 · 0 citations · 53 references
Medicine

Abstract

Reaction optimization is a very time- and resource-intensive process, as optimal reaction conditions depend highly on specific electronic and steric properties of the substrate identity and require extensive fine-tuning of synthetic conditions to arrive at the highest-yielding conversions. Amide couplings, which comprise nearly 40% of synthetic transformations performed in a medicinal chemistry setting, present a particularly challenging context for predictive modeling given the diversity of coupling agents and reaction parameters that must be matched to substrate reactivity. Amide coupling reaction data is particularly well-suited for machine-learning approaches that predict optimal reaction conditions given a particular substrate feature set. Herein, we report a platform for standardizing and filtering open-source reaction data from the ORD (Open Reaction Database) and using this machine-readable data set of 3800 amide coupling reactions to evaluate 13 machine learning models. These include linear, tree-based, kernel method, instance-based, neural network, and ensemble architectures in the yield prediction and classification of coupling agents in amide coupling reactions. Yield prediction remained a difficult task because of the complexity of our reaction data, achieving R 2 scores of only 0.61. However, the models were largely successful in classifying reactions to their ideal coupling agent category (carbodiimide-based, uronium salts, or phosphonium salts) with ensemble- and kernel-based models achieving up to 87% accuracy. To further validate this approach, we deployed the classification models on literature-reported data not in the ORD database, achieving similarly high predictive performance. Our results demonstrate that kernel methods and ensemble-based architectures perform significantly better than other models such as linear or single-tree. Additionally, molecular environment features, captured by XYZ coordinates, three-dimensional features, and Morgan fingerprints around reactive functional groups, boosted model predictivity more than bulk material properties derived from SMILES such as molecular weight and log P.

Read PDF