Accurate modeling of protein-peptide interactions is essential for understanding fundamental biological processes and designing peptide-based drugs. However, predicting the complex structures of these interactions remains challenging, primarily due to the high conformational flexibility of peptides. To support a fair and systematic evaluation of recent deep learning (DL) approaches, we introduce PepPCBench, a benchmarking framework tailored to assess protein folding neural networks (PFNNs) in protein-peptide complex prediction. As part of this framework, we curated PepPCSet, a data set of 261 experimentally resolved complexes with peptides ranging from 5 to 30 residues. We benchmark five full-atom PFNNs, including AlphaFold3 (AF3), AlphaFold-Multimer (AFM), Chai-1, HelixFold3 (HF3), and RoseTTAFold-All-Atom (RFAA), using comprehensive evaluation metrics. Our benchmarking reveals meaningful performance differences among these methods and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy. While AF3 shows strong performance in structure prediction, further analysis indicates that confidence metrics correlate poorly with experimental binding affinities, underscoring the need for improved scoring strategies and generalizability. By providing a reproducible and extensible framework, PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction.
Silong Zhai, Huifeng Zhao, Jike Wang et al.· Journal of Chemical Informat...· 13 citations· ⚡1
Abstract Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein–cyclic peptide (P–cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson–Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P–cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P–cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P–cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P–cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P–cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of −0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = −0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P–cp interactions.
The integration of organic synthesis with enzymatic catalysis offers a promising route toward efficient and sustainable construction of complex molecules. While organic synthesis enables diverse transformations, enzymatic catalysis enhances stereoselectivity under mild conditions, improving cost-effectiveness and environmental impact. However, current enzymatic synthesis planning algorithms face challenges in formulating robust hybrid organic–enzymatic strategies. Key issues include the difficulty in devising hybrid planning approaches and the reliance on template-based enzyme recommendations, which limits their adaptability across diverse scenarios. Here we show ChemEnzyRetroPlanner, an open-source hybrid synthesis planning platform that combines organic and enzymatic strategies with AI-driven decision-making. The platform features advanced computational modules, including hybrid retrosynthesis planning, reaction condition prediction, plausibility evaluation, enzymatic reaction identification, enzyme recommendation, and in silico validation of enzyme active sites. A central innovation is the RetroRollout* search algorithm, which outperforms existing tools in planning synthesis routes for organic compounds and natural products across multiple datasets. ChemEnzyRetroPlanner provides an intuitive graphical interface and programmatic APIs for scalability, while leveraging the chain-of-thought strategy and the Llama3.1 model to autonomously activate hybrid synthesis strategies for diverse scenarios. The results indicate that this fully automated, open-source system holds potential value for improving the efficiency and sustainability of molecular synthesis. The integration of organic and enzymatic synthesis enhances molecule construction efficiency. Here, the authors present ChemEnzyRetroPlanner, an AI-driven platform for automated hybrid synthesis planning, improving synthesis route efficiency and sustainability.
The integration of large language models (LLMs) into drug design is gaining momentum; however, existing approaches often struggle to effectively incorporate three-dimensional molecular structures. Here, we present Token-Mol, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens. Built on a transformer decoder and trained with causal masking, Token-Mol introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream applications. The model surpasses existing methods, improving molecular conformation generation by over 10% and 20% across two datasets, while outperforming token-only models by 30% in property prediction. In pocket-based molecular generation, it enhances drug-likeness and synthetic accessibility by approximately 11% and 14%, respectively. Notably, Token-Mol operates 35 times faster than expert diffusion models. In real-world validation, it improves success rates and, when combined with reinforcement learning, further optimizes affinity and drug-likeness, advancing AI-driven drug discovery. In this work the authors present Token-Mol, a token-only 3D drug design model, which deploys the Gaussian cross-entropy (GCE) loss function for regression tasks. It exhibits superior performance in molecular conformation generation, property prediction, and pocket-based generation, thus opening up new avenues for drug design.
Jike Wang, Rui Qin, Mingyang Wang et al.· Nature Communications· 30 citations· ⚡1
The discovery of CAP-Gly domain-containing linker protein 1(CLIP1)-Leukocyte tyrosine kinase (LTK) as an oncogenic fusion reveals a unique dependency not only on LTK kinase activity but also on CLIP1-mediated multimerization, a noncatalytic function that drives oncogenic signaling. While this fusion is currently targeted with anaplastic lymphoma kinase inhibitors, their exclusive focus on kinase inhibition leaves the scaffolding function intact, necessitating a complete protein clearance strategy. Here, we report the AI-guided development of a first-in-class proteolysis-targeting chimera (PROTAC) designed to selectively degrade the CLIP1-LTK fusion protein. By integrating deep learning models for ternary complex prediction with structure-based molecular optimization, we designed DCL05, an orally bioavailable degrader of CLIP1-LTK fusion protein, achieving picomolar degradation potency (DC50 = 40 pM) and robust antitumor activity. DCL05 consistently outperformed existing kinase inhibitors across a broad spectrum of LTK resistance-associated mutations, both in vitro and in vivo. Collectively, our study explores resistance-associated contexts of LTK and establishes a structure-guided PROTAC development pipeline, providing a promising therapeutic strategy for overcoming acquired resistance in kinase-driven cancers.
Shicheng Chen, Haiting Duan, S. Zhong et al.· Proceedings of the National...· 0 citations
Enzymatic reactions play an emerging role in a broad spectrum of scientific and industrial applications. The inherent complexity of enzymes, such as their substrate specificity, conformational flexibility, and the vast diversity of reactions involved, poses substantial challenges for the advanced computational prediction of enzymatic reactions with desirable accuracy. Moreover, existing approaches are mostly tailored for a specific sub-task, such as substrate prediction or binding site annotation, which limits their applicability. In this study, we introduce ERAM, a task-agnostic multimodal learning framework capable of addressing a broad range of downstream applications with both accuracy and efficiency. ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data. In enzyme retrieval tasks, ERAM achieves an improvement of 28.31% in mean average precision compared with the state-of-the-art (SOTA) method, CREEP. In substrate prediction tasks, ERAM outperforms the SOTA method ESP, achieving average improvements of 35.53% and 22.97% in Matthews correlation coefficient across two datasets. Additionally, ERAM exhibits commendable interpretability by assigning higher attention weights to binding sites, resulting in lower false-positive rates (42.36%) and higher overlap scores (70.59%) in the unsupervised binding site prediction task compared to RXNAA Mapper. By learning embeddings of substrates, enzymes, and products within a unified knowledge graph latent space, ERAM demonstrates its potential as a versatile and effective tool for enzyme catalysis research.
A novel committor learning framework grounded in the AlphaFold 3 paradigm is proposed that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.
Jintu Zhang, Zichang Jin, Huifeng Zhao et al.· 0 citations