Skip to content

Category

quantum computing

321 papers

#computer vision Nov 2025

Improving the predictive performance of binding affinities and poses for protein–cyclic peptide complexes through fine-tuned MM/PBSA(GBSA)-based methods

Abstract Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein–cyclic peptide (P–cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson–Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P–cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P–cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P–cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P–cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P–cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of −0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = −0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P–cp interactions.

Huifeng Zhao, Jianxiang Huang, Gaoqi Weng et al. · 8 citations

Revisiting Protein-Protein Docking: A Systematic Evaluation Framework

Protein-protein interactions play pivotal roles in a wide range of biological processes. Determining the atomic-level structures of protein-protein complexes is indispensable for elucidating macromolecular interaction mechanisms and advancing structure-based drug design. Protein-protein docking, as one of the leading computational approaches for predicting complex structures, has seen considerable progress but requires rigorous evaluation in practical applications. In this study, we proposed a comprehensive benchmarking framework to evaluate 11 docking methods spanning traditional (HDOCK, PatchDock, PIPER, ZDOCK) and deep learning (DL)-based (EquiDock, ElliDock, EBMDock, GeoDock, DiffDock-PP, AlphaFold-Multimer, AlphaFold3) approaches. Our framework incorporates the classical DockingBenchmark 5.5 data set for evaluating flexible docking, introduces a newly curated data set (AACBench) for antibody-antigen complex docking, and establishes the PPCBench data set to examine the out-of-distribution (OOD) generalization capabilities of DL-based methods. In docking against apo structures, AlphaFold3 achieves a superior top-5 success rate of 77.98%, whereas the traditional approach HDOCK reaches merely 12.84%, despite its highest top-5 success rate of 85.24% when docking against holo structures. For antibody-antigen docking, AlphaFold3 remains the most accurate method (top-5 success rate: 31.78%) and substantially outperforms AlphaFold-Multimer in modeling the CDR-H3 loop. In OOD generalization tests, all DL-based models exhibit markedly reduced performance on the PPCBench data set. Overall, our work establishes a unified benchmarking framework that enables systematic evaluation of docking methods across diverse tasks and provides critical insights into the strengths and limitations of current docking strategies, thereby informing future developments in protein-protein docking research.

Linlong Jiang, Ke Zhang, Kai Zhu et al. · 3 citations

Macrocycle-DB: a comprehensive database for macrocycle-based drug discovery

Abstract Macrocycles have gained significant attention in drug design owing to their distinctive structural and physicochemical features. Despite the abundance of available experimental data, there remains a need for a centralized resource to support macrocycle-based drug discovery. Here, we present Macrocycle-DB, the most extensive online database dedicated to macrocycles, featuring 45 525 compounds, including 76 approved drugs and 105 clinical candidates that target 2533 proteins. The database offers comprehensive structural information, experimental bioactivity data, and physicochemical properties for each macrocycle, along with co-crystal structures to visualize protein–ligand interactions. Additionally, Macrocycle-DB provides specialized descriptors, scaffold and linker details for synthetic macrocycles, and high-quality downloadable datasets to facilitate computational drug design. Macrocycle-DB is freely accessible at https://macro-db.dpbio.tech/ and https://macro-db.cn/.

Minchuan Jiang, Tianyu Liu, Muzammal Hussain et al. · 6 citations

Overcoming Resistance in the Androgen Receptor: Rational and Strategic Design of Advanced Antagonists.

ConspectusProstate cancer (PCa) is the most prevalent malignancy among men worldwide, with its pathogenesis and progression heavily reliant on the sustained activation of the androgen receptor (AR) signaling pathway. The AR, a transcription factor of nuclear receptor superfamily, serves as the most privileged therapeutic target in PCa, as evidenced by the clinical efficacy of first- and second-generation AR antagonists. Current clinically available AR antagonists exclusively target the ligand binding pocket (LBP), suppressing tumor proliferation through competitive inhibition of androgen binding and subsequent blockade of AR signaling transduction. However, their therapeutic utility is invariably limited by acquired resistance mechanisms, including point mutations that alter LBP specificity, AR gene amplification leading to receptor overexpression, and the emergence of constitutively active splice variants that bypass ligand-dependent activation. Thus, the development of novel AR antagonists featuring innovative mechanisms and structural scaffolds is imperative to overcome resistance to antiandrogen therapy. However, the AR exhibits significant structural flexibility, and the lack of antagonist-bound crystal structures has hindered structure-based rational drug design. In this Article, we summarize our advances in elucidating the molecular mechanisms underlying AR conformational regulation and highlight our progress in the structure-based design and development of novel AR antagonists. First, our molecular dynamic (MD) studies collectively elucidate the molecular mechanisms by which the AR ligand binding domain (LBD) regulates its functional states through dynamic conformational changes mediated by distinct allosteric pathways when bound to agonists or antagonists, providing atomic-level insights and structural basis for drug development. Then, we successfully identified structurally diverse lead compounds targeting the LBP through various integrated approaches combining MD simulations, structure-based virtual screening (SBVS), and systematic biological evaluation. These compounds exhibited potent activity against clinically relevant AR mutations F877L, W742C, T878A, and H875Y, demonstrating their potential to overcome mutations-driven resistance. Further, we explored non-LBP mediated strategies for AR antagonism, including: (1) targeting the allosteric binding sites on LBD; (2) identification of novel druggable binding sites; and (3) targeting alternative domains beyond the LBD. As a paradigm-shifting example, we proposed inhibition of AR LBD dimerization as a novel mechanism of action for LBP-targeting AR antagonists. Building upon this insight, we characterized a promising pocket at the dimer interface, designated the Dimerization Interface Pocket (DIP), and developed first-in-class antagonists specifically targeting this site, which exhibit exceptional therapeutic potential. Collectively, these multipronged strategies not only highlight the power of computation-driven approaches in drug discovery but also yield a diverse pipeline of resistance-targeting candidates, directly addressing the unmet clinical need in advanced PCa.

Xin Chai, Tingjun Hou, Dan Li · 0 citations
#computer vision Jan 2026

NavDB: A Comprehensive Database for Voltage-Gated Sodium Channels Modulators and Targets

Voltage-gated sodium channels (VGSCs/Navs) are essential targets for the treatment of numerous neurological, muscular, and cardiac disorders. Despite the increasing clinical interest in subtype-selective modulators, current public databases provide fragmented and inconsistent information on VGSC-related compounds and targets, particularly lacking coverage on peptides. To address this limitation, we developed NavDB, a specialized and open-access database focusing on VGSC modulators and targets. NavDB integrates 8023 curated data records covering 5168 compounds, including small molecules, toxins, drugs, and peptides, along with comprehensive annotations on biological activity, druggability, and structural feature. NavDB also features advanced functions such as text-based and structure-based search, peptide similarity matching, and AI-powered property prediction. Moreover, the database offers high-quality 3D visualizations of targets and peptides, with disulfide bond and signal peptide annotations. All data are freely downloadable to support both experimental and computational drug discovery. NavDB is publicly available at: http://cadd.zju.edu.cn/navdb/.

Gaoang Wang, Jiahui Yu, Haiyi Chen et al. · 0 citations

STE-DC2I Uncovers Driver Genes in Colorectal Cancer Subtypes Using Symbolic Trajectory-Embedded Dark Causal Inference

Colorectal cancer (CRC) exhibits substantial molecular heterogeneity, necessitating the inference of subtype-specific driver genes and their interactions for drug-target discovery and precision oncology. Prior studies often fail to capture subtle, latent nonlinear regulatory mechanisms (dark causal relationships) driving tumor progression in specific subtypes. Here, we develop an explainable intelligence computational framework, Symbolic Trajectory-Embedded Dark Causal Interaction Inference (STE-DC2I), which combines symbolic trajectory embedding with historical prediction mechanisms to model nonmonotonic oscillatory dependencies between genes. Integrating single-cell transcriptomic and multiomics profiles from malignant epithelial subpopulations, STE-DC2I classifies CRC subtypes, reconstructs developmental trajectories, and uncovers interpretable subtype-specific driver genes with functional relevance. Unlike correlation-based and explicit causal approaches, STE-DC2I captures weak yet biologically critical regulatory signals, outperforming state-of-the-art methods in predicting subtype-specific CRC driver genes. Functional assays in CRC cell lines (in vitro) validated nine predicted driver genes, highlighting their therapeutic potential.This work systematically explores dark causal interactions between genes in CRC subtypes. STE-DC2I offers interpretable insights and a generalizable strategy for CRC drug-target discovery.

Meng Huang, Huijin Hu, Ming Li et al. · 0 citations

How to efficiently characterize the interaction pathways of protein-ligand recognition? A comparative analysis on enhanced sampling approaches.

It is evidenced that many elaborately designed molecules that can interact well with the binding pocket of their target fail to exhibit activity in wet-lab experiments. This may associate with the interacting process of drug-target recognition. To efficiently characterize the drug-target interacting process, various enhanced sampling technologies have been proposed; yet, very few studies have systemically investigated whether the settings of these simulations are favorable to characterize the purposed tasks. Here, by comparing two popular enhanced sampling technologies, namely, the well-temped metadynamics and random acceleration molecular dynamics (RAMD), we systemically investigate the strategies to efficiently characterize the dissociating process of protein-ligand interactions. Two target families are employed for the analysis, including the kinase family (represented by TRK1) that represents the interaction-pathway obvious systems and the nuclear receptor family (represented by THRβ) that represents the interaction-pathway unobvious systems. Our results suggest that (1) in terms of maintaining stability of the protein structure, MetaD at various simulation conditions and RAMD with a large random force are good choice; (2) drug residence time derived from both MetaD and RAMD based on various parameters shows reasonable correlation to the experimental binding strength of the ligands, but RAMD usually runs with much less simulation time; and (3) both enhanced sampling methods result in reasonably consistent pathway preference for the two target families. Taken together, it will be much time-saving to utilize RAMD with high random force for interaction pathway exploration for both the pathway obvious and unobvious systems if the protein keeps stable in the simulation; otherwise, MetaD with a high bias factor is proposed to balance the computational accuracy and efficiency for the exploration.

Zhiliang Jiang, Mingyun Shen, Zhe Wang et al. · 1 citation
#computer vision Jan 2026

Understanding the Kinetic Mechanism of Ligands Stabilizing the RAS-CYPA Interaction

Molecular glues, including protein degraders and protein-protein interaction (PPI) stabilizers, have emerged as a new paradigm of drug design for regulating interactions between biomacromolecules; yet it is still a challenge for rational design of molecular glues. KRAS, as a prevalent oncogenic driver, is notoriously difficult to target by traditional small molecular drugs due to its challenging binding surface and frequent mutations. Although the small molecular drug RMC7977 has been designed as a PPI stabilizer for stabilizing the inherently weak RAS-CYPA interaction, the precise molecular mechanism underlying its stabilization effect and selectivity difference requires a deeper understanding. To this end, we leverage an integrated computational strategy combining molecular dynamics (MD) simulation, end-point binding free-energy calculation, and enhanced sampling technologies to elucidate the dynamic characteristics of RAS-ligand-CYPA interactions. Our result exhibits a high correlation between the predicted binding affinities and the experimental observations, demonstrating that RMC7977, acting as a strong PPI stabilizer, significantly enhances the stability of the KRAS-CYPA interaction, where, by delicately remodeling the protein-protein interface, the drug optimizes various interactions. Moreover, the results also uncover the dynamic process of stabilizer-mediated KRAS-CYPA stabilization and the mechanistic origin of the binding selectivity. This study provides essential molecular-level insights into RMC7977's function and offers a valuable computational framework for evaluating the stabilization effect of ligands targeting the KRAS-CYPA and other challenging PPI systems.

Kexin Xu, Mingyun Shen, Zhe Wang et al. · 0 citations

Computational and AI-Driven Ecosystem for Structure-Based Covalent Drug Discovery.

ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.

Shi Li, Hongyan Du, Xujun Zhang et al. · 4 citations

DRHIN: An Integrated and Interactive Web Server for Drug Repositioning

Drug repositioning (DR) identifies new therapeutic uses for approved drugs, reducing development burdens and offering safer treatment options for patients. While high-throughput technologies generate complex, large-scale multiomics data, existing DR tools struggle to comprehensively analyze the resulting biological networks. To address this challenge, we present DRHIN, an integrated, interactive web server for DR over heterogeneous information networks (HINs) using advanced deep learning techniques. DRHIN integrates transcriptomics, proteomics, and microbiome data, incorporating eight biological entities and 19 association types to build diverse HINs and elucidate the underlying molecular mechanisms. It includes 19 state-of-the-art graph representation algorithms, enabling flexible training, comparison, and evaluation of heterogeneous network data. The platform provides a code-free portal supporting three key predictive tasks: discovering drug-disease associations, repurposing existing drugs for new indications, and identifying potential therapies for specific diseases, making analyses accessible and reproducible. Leveraging high-performance computing, DRHIN efficiently processes million-scale networks, ensuring practical applicability in real-world scenarios. The web server is freely accessible at http://drhin.tianshanzw.cn.

Bowei Zhao, Dongxu Li, Yue Yang et al. · 6 citations · ⚡1
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

Structure-based machine learning algorithms have been utilized to predict the properties of protein-protein interaction (PPI) complexes, such as binding affinity, which is critical for understanding biological mechanisms and disease treatments. While most existing algorithms represent PPI complex graph structures at the atom-scale or residue-scale, these representations can be computationally expensive or may not sufficiently integrate finer chemical-plausible interaction details for improving predictions. Here, we introduce MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently. This framework maps proteins onto a concise CG-scale complex graph, where nodes represent CG beads and edges encode chemically plausible interactions. The GNN-based encoder is tailored to extract high-quality representations from this graph, efficiently capturing the overall properties of the protein complex structure. Extensive experiments on three different downstream PPI property prediction tasks demonstrate that MCGLPPI achieves competitive performance compared with the counterparts at the atom- and residue-scale, but with only a third of the computational resource consumption. Furthermore, the CG-scale pre-training on protein domain-domain interaction structures enhances its predictive capabilities for PPI tasks. MCGLPPI offers an effective and efficient solution for PPI overall property predictions, serving as a promising tool for the large-scale analysis of biomolecular interactions.

Yang Yue, Shu Li, Yihua Cheng et al. · 14 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.