Jul 2026· Journal of Chemical Information and Modeling· Vol 66 15, pp.
8908-8922
· 0 citations· 39 references
Medicine
TL;DR
This work introduces a novel approach to transform the concept of "how to design chemical analogs like a medicinal chemist" into a primary training objective for generative models by focusing on matched molecular pair transformations (MMPTs) as the fundamental unit of chemical modifications.
Abstract
Chemical analog design in the hit-to-lead and lead optimization stages of drug discovery relies on systematic structural modifications, often guided by medicinal chemistry intuition. Although a matched molecular pair (MMP) provides an interpretable framework to capture intuition, models trained on individual MMP instances face significant limitations, such as bias toward frequent transformations in historical data. This work introduces a novel approach to transform the concept of "how to design chemical analogs like a medicinal chemist" into a primary training objective for generative models by focusing on matched molecular pair transformations (MMPTs) as the fundamental unit of chemical modifications. This allows for a more generalizable and context-independent representation of medicinal chemistry intuition, enabling the application of the same transformation priors across different projects, regardless of the target or indication. Using a consistently curated ChEMBL-derived data set, we compared a transformation-centric foundation model (MMPT-FM) with multiple MMP-based generative formulations trained on the same underlying data. Furthermore, performance is assessed through challenging within-patent and cross-patent real-world test cases derived from drug discovery patents. The MMPT-FM model achieves comparable or improved recall metrics across all test cases, demonstrating particularly strong performance for low-frequency and previously unseen transformations. This work not only shifts the paradigm of learning and utilizes medicinal chemistry intuition in an efficient and scalable manner in the AI for drug discovery era but also establishes a significant competitive advantage by enabling the drug discovery industry to encode decades of collective medicinal chemistry knowledge─both public and proprietary─into a scalable foundation generative model that could help researchers design chemical analogs.
A critical perspective is provided on how generative models are shaping the future of rational and reliable drug design, including automated synthesis planning, retrosynthesis prediction, and multi‐objective optimization.
Rania Ehab Koshty, Manar Ahmed Shehata, Ahmed M. Gab Allah et al.· ChemistrySelect· 0 citations
ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.
Veincent Yap, Pan Xu, Frankie S. Mak et al.· Journal of Cheminformatics· 0 citations
Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.
Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al.· 1 citation
Packora is presented, a flow-based generative model for molecular CSP that jointly predicts atomic coordinates and the lattice from molecular graphs that outperforms the baselines on both structure generation and ranking benchmarks.
Nayoung Kim, Kiyoung Seong, Sungsoo Ahn· 0 citations
This work introduces Sample Efficient Generative Optimization (SEGO), a framework for Bayesian optimization on adaptively generated molecules, and attains state-of-the-art performance on the practical molecular optimization (PMO) benchmark using only one tenth of the oracle calls consumed by other methods.
S. Kopf, Cristina Nevado, P. Schwaller· 0 citations
Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.
Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al.· 0 citations