Skip to content

Category

large language models

490 papers

#computer vision May 2025

A Unified Deep Graph Model for Identifying the Molecular Categories of Ligands Targeting Nuclear Receptors

To fulfill functions for differentially regulating the downstream signaling pathways, functional ligands (i.e., agonists or antagonists) targeting nuclear receptors (NRs) are designed to stabilize different conformations (active or inactive) of the proteins. However, in practical applications, it is usually difficult to determine the molecular category of an NR ligand because these molecules all bind in the same location of an NR protein, namely, the ligand-binding pocket (LBP). Considering that ligands with different properties (agonists or antagonists) prefer to bind with differential conformations of NRs, it is possible to identify the molecular type of a given ligand through the differential binding environment (active or inactive conformations) of the protein-ligand interaction. Therefore, in this study, we established a unified model (NRIGN) based on the deep graphic architecture to discriminate agonists and antagonists targeting 26 successful or in-clinical-trial NR targets. Our result shows that NRIGN achieves an excellent prediction accuracy (ACC >0.95) and is robust enough to be applied in various real-world scenarios, such as predicting the molecular type of ligands in crystallized NR structures, ligands with multiple NR activities, and ligands with their types altered by target mutations. The proposed model is expected to promote rational design of drugs targeting NR proteins.

Kaimo Yang, Dejun Jiang, Qirui Deng et al. · 2 citations

Effective generation of heavy-atom-free triplet photosensitizers containing multiple intersystem crossing mechanisms based on deep learning

Photodynamic therapy (PDT) is a clinically approved therapeutic modality that has demonstrated significant potential for cancer treatment, and triplet photosensitizers (PSs) play a key role in its efficacy. Despite deep learning having emerged as a next-generation tool for material discovery, existing methods mainly target a limited subset of triplet PSs, such as thermally activated delayed fluorescence (TADF) materials, neglecting the critical intersystem crossing (ISC) between the high-lying singlet and triplet states (ΔESnTn). To overcome this limitation, we compiled a comprehensive dataset (∼1.90 × 109) of triplet PSs encompassing various ISC mechanisms. Then, we proposed a novel strategy that incorporates two models: a fragment-based model (Frag-MD) and a character-based model (MD), both integrating a conditional transformer, recurrent neural networks, and reinforcement learning. In silico experiments revealed that the Frag-MD model outperforms the MD model in generating larger conjugated motifs with higher average ring numbers and atom counts; while the MD model generates twice as many unique motifs and excels in novelty and diversity, as evaluated by conditional and MOSES metrics. Therefore, our approach is highly effective for modifying conjugated motifs and designing novel triplet PSs. Notably, the recently reported high-efficiency triplet PSs have been re-identified through ablation experiments using our proposed models, which target ΔESnTn and significantly outperform traditional baselines, achieving a prediction accuracy of 73% versus 4%. Our approach holds the potential to establish a new paradigm for discovering novel PSs applicable in PDT.

Kepeng Chen, Xiaoting Zhang, Jike Wang et al. · 3 citations

Revisiting Protein-Protein Docking: A Systematic Evaluation Framework

Protein-protein interactions play pivotal roles in a wide range of biological processes. Determining the atomic-level structures of protein-protein complexes is indispensable for elucidating macromolecular interaction mechanisms and advancing structure-based drug design. Protein-protein docking, as one of the leading computational approaches for predicting complex structures, has seen considerable progress but requires rigorous evaluation in practical applications. In this study, we proposed a comprehensive benchmarking framework to evaluate 11 docking methods spanning traditional (HDOCK, PatchDock, PIPER, ZDOCK) and deep learning (DL)-based (EquiDock, ElliDock, EBMDock, GeoDock, DiffDock-PP, AlphaFold-Multimer, AlphaFold3) approaches. Our framework incorporates the classical DockingBenchmark 5.5 data set for evaluating flexible docking, introduces a newly curated data set (AACBench) for antibody-antigen complex docking, and establishes the PPCBench data set to examine the out-of-distribution (OOD) generalization capabilities of DL-based methods. In docking against apo structures, AlphaFold3 achieves a superior top-5 success rate of 77.98%, whereas the traditional approach HDOCK reaches merely 12.84%, despite its highest top-5 success rate of 85.24% when docking against holo structures. For antibody-antigen docking, AlphaFold3 remains the most accurate method (top-5 success rate: 31.78%) and substantially outperforms AlphaFold-Multimer in modeling the CDR-H3 loop. In OOD generalization tests, all DL-based models exhibit markedly reduced performance on the PPCBench data set. Overall, our work establishes a unified benchmarking framework that enables systematic evaluation of docking methods across diverse tasks and provides critical insights into the strengths and limitations of current docking strategies, thereby informing future developments in protein-protein docking research.

Linlong Jiang, Ke Zhang, Kai Zhu et al. · 3 citations

MetalloDock: Decoding Metalloprotein-Ligand Interactions via Physics-Aware Deep Learning for Metalloprotein Drug Discovery.

Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDock, the first DL-based docking framework specially designed for metalloprotein targets. By innovatively integrating an autoregressive spatial decoding engine with a physics-constrained geometric generation paradigm, MetalloDock can precisely reconstruct metal coordination geometries and accurately capture metal-ligand interactions, which enhance both the accuracy of metalloprotein-ligand docking and binding affinity prediction. Extensive evaluations on our custom-built benchmark data set demonstrate that MetalloDock outperforms existing methods, including AlphaFold3, in docking success rate and virtual screening performance for metalloprotein targets. In real-world applications, MetalloDock successfully identified multiple novel hit compounds in a virtual screening campaign targeting the prostate-specific membrane antigen. Additionally, it enabled rational drug design for acidic polymerase endonuclease, leading to the discovery of potent inhibitors. These results highlight the broad applicability of MetalloDock in accelerating metalloprotein-targeted drug discovery and provide a standardized framework for future evaluation of metalloprotein-specific docking algorithms.

Hui Zhang, Xujun Zhang, Qun Su et al. · 5 citations

STE-DC2I Uncovers Driver Genes in Colorectal Cancer Subtypes Using Symbolic Trajectory-Embedded Dark Causal Inference

Colorectal cancer (CRC) exhibits substantial molecular heterogeneity, necessitating the inference of subtype-specific driver genes and their interactions for drug-target discovery and precision oncology. Prior studies often fail to capture subtle, latent nonlinear regulatory mechanisms (dark causal relationships) driving tumor progression in specific subtypes. Here, we develop an explainable intelligence computational framework, Symbolic Trajectory-Embedded Dark Causal Interaction Inference (STE-DC2I), which combines symbolic trajectory embedding with historical prediction mechanisms to model nonmonotonic oscillatory dependencies between genes. Integrating single-cell transcriptomic and multiomics profiles from malignant epithelial subpopulations, STE-DC2I classifies CRC subtypes, reconstructs developmental trajectories, and uncovers interpretable subtype-specific driver genes with functional relevance. Unlike correlation-based and explicit causal approaches, STE-DC2I captures weak yet biologically critical regulatory signals, outperforming state-of-the-art methods in predicting subtype-specific CRC driver genes. Functional assays in CRC cell lines (in vitro) validated nine predicted driver genes, highlighting their therapeutic potential.This work systematically explores dark causal interactions between genes in CRC subtypes. STE-DC2I offers interpretable insights and a generalizable strategy for CRC drug-target discovery.

Meng Huang, Huijin Hu, Ming Li et al. · 0 citations

Discovery of Novel Nonsteroidal SGRMs of Sulfonamide-2-Oxo-Tetrahydroquinoline Derivatives by Carbonyl Migration.

Glucocorticoids (GCs) are limited by severe side effects, driving the development of selective glucocorticoid receptor modulators (SGRMs) with improved therapeutic profiles. We previously development the SGRM lead B53, which suffered from poor metabolic stability. In this study, structure-guided optimization of B53 yielded 43 novel sulfonamide derivatives. Among them, D8, which contained 2-oxo-tetrahydroquinoline by carbonyl migration form B53, manifests an excellent SGRM with remarkable transrepression potency (IC50NF-κB = 0.9 nM) superior to dexamethasone (IC50 NF-κB = 5.0 nM). Besides, D8 exhibits a significantly higher specificity for GR over AR, MR, and PR and exhibited less adverse effects on osteoprotegerin. Furthermore, D8 demonstrated improved metabolic stability and optimized binding mode within the GR LBD. In vivo, oral administration of D8 significantly alleviated dermatitis and autoimmune hepatitis in mouse models, underscoring its therapeutic potential and validating our design strategy.

Xiaodong Bao, Yuxin Zhou, Zhaoxu Yang et al. · 1 citation

How to efficiently characterize the interaction pathways of protein-ligand recognition? A comparative analysis on enhanced sampling approaches.

It is evidenced that many elaborately designed molecules that can interact well with the binding pocket of their target fail to exhibit activity in wet-lab experiments. This may associate with the interacting process of drug-target recognition. To efficiently characterize the drug-target interacting process, various enhanced sampling technologies have been proposed; yet, very few studies have systemically investigated whether the settings of these simulations are favorable to characterize the purposed tasks. Here, by comparing two popular enhanced sampling technologies, namely, the well-temped metadynamics and random acceleration molecular dynamics (RAMD), we systemically investigate the strategies to efficiently characterize the dissociating process of protein-ligand interactions. Two target families are employed for the analysis, including the kinase family (represented by TRK1) that represents the interaction-pathway obvious systems and the nuclear receptor family (represented by THRβ) that represents the interaction-pathway unobvious systems. Our results suggest that (1) in terms of maintaining stability of the protein structure, MetaD at various simulation conditions and RAMD with a large random force are good choice; (2) drug residence time derived from both MetaD and RAMD based on various parameters shows reasonable correlation to the experimental binding strength of the ligands, but RAMD usually runs with much less simulation time; and (3) both enhanced sampling methods result in reasonably consistent pathway preference for the two target families. Taken together, it will be much time-saving to utilize RAMD with high random force for interaction pathway exploration for both the pathway obvious and unobvious systems if the protein keeps stable in the simulation; otherwise, MetaD with a high bias factor is proposed to balance the computational accuracy and efficiency for the exploration.

Zhiliang Jiang, Mingyun Shen, Zhe Wang et al. · 1 citation
#computer vision Jan 2026

Understanding the Kinetic Mechanism of Ligands Stabilizing the RAS-CYPA Interaction

Molecular glues, including protein degraders and protein-protein interaction (PPI) stabilizers, have emerged as a new paradigm of drug design for regulating interactions between biomacromolecules; yet it is still a challenge for rational design of molecular glues. KRAS, as a prevalent oncogenic driver, is notoriously difficult to target by traditional small molecular drugs due to its challenging binding surface and frequent mutations. Although the small molecular drug RMC7977 has been designed as a PPI stabilizer for stabilizing the inherently weak RAS-CYPA interaction, the precise molecular mechanism underlying its stabilization effect and selectivity difference requires a deeper understanding. To this end, we leverage an integrated computational strategy combining molecular dynamics (MD) simulation, end-point binding free-energy calculation, and enhanced sampling technologies to elucidate the dynamic characteristics of RAS-ligand-CYPA interactions. Our result exhibits a high correlation between the predicted binding affinities and the experimental observations, demonstrating that RMC7977, acting as a strong PPI stabilizer, significantly enhances the stability of the KRAS-CYPA interaction, where, by delicately remodeling the protein-protein interface, the drug optimizes various interactions. Moreover, the results also uncover the dynamic process of stabilizer-mediated KRAS-CYPA stabilization and the mechanistic origin of the binding selectivity. This study provides essential molecular-level insights into RMC7977's function and offers a valuable computational framework for evaluating the stabilization effect of ligands targeting the KRAS-CYPA and other challenging PPI systems.

Kexin Xu, Mingyun Shen, Zhe Wang et al. · 0 citations

Computational and AI-Driven Ecosystem for Structure-Based Covalent Drug Discovery.

ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.

Shi Li, Hongyan Du, Xujun Zhang et al. · 4 citations

DRHIN: An Integrated and Interactive Web Server for Drug Repositioning

Drug repositioning (DR) identifies new therapeutic uses for approved drugs, reducing development burdens and offering safer treatment options for patients. While high-throughput technologies generate complex, large-scale multiomics data, existing DR tools struggle to comprehensively analyze the resulting biological networks. To address this challenge, we present DRHIN, an integrated, interactive web server for DR over heterogeneous information networks (HINs) using advanced deep learning techniques. DRHIN integrates transcriptomics, proteomics, and microbiome data, incorporating eight biological entities and 19 association types to build diverse HINs and elucidate the underlying molecular mechanisms. It includes 19 state-of-the-art graph representation algorithms, enabling flexible training, comparison, and evaluation of heterogeneous network data. The platform provides a code-free portal supporting three key predictive tasks: discovering drug-disease associations, repurposing existing drugs for new indications, and identifying potential therapies for specific diseases, making analyses accessible and reproducible. Leveraging high-performance computing, DRHIN efficiently processes million-scale networks, ensuring practical applicability in real-world scenarios. The web server is freely accessible at http://drhin.tianshanzw.cn.

Bowei Zhao, Dongxu Li, Yue Yang et al. · 6 citations · ⚡1

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.