Skip to content

Category

machine learning

3,595 papers

#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

The authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks, which signifies a substantial leap forward in mRNA research and therapeutic development.

Ying Xiong, Aowen Wang, Yu Kang et al. · 23 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations
#machine learning Open access May 2025

Token-Mol 1.0: tokenized drug design with large language models

Token-Mol is presented, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens, which introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream applications.

Ji-Ke Wang, Rui Qin, Mingyang Wang et al. · 30 citations · ⚡1
#machine learning Open access Nov 2025

A fused deep learning approach to transform drug repositioning

Drug repositioning holds promise for discovering new therapeutic applications for existing drugs, accelerating drug development and reducing associated costs. However, current methodologies encounter difficulties in managing diverse network representations, tackling cold start issues, and handling intrinsic attribute representations. Here we introduce a Unified Knowledge-Enhanced deep learning framework for Drug Repositioning (UKEDR), which integrates knowledge graph embedding, pre-training strategies, and recommendation systems to address these challenges. To overcome the cold start issue, UKEDR utilizes a semantic similarity-driven embedding approach. Our evaluations show that UKEDR performs better than various baselines, including classical machine learning, network-based, and deep learning approaches. In cold start scenarios, it demonstrates an improved capability in handling unseen nodes and generalizing to new compounds. The model also demonstrates strong robustness on imbalanced datasets and shows excellent generalization capabilities in specific drug-centric and disease-centric cold-start scenarios, validating its potential for real-world applications. Drug repositioning offers a promising avenue for accelerating drug development, yet existing methods struggle with network diversity, cold start issues, and intrinsic attribute representation. Here, the authors introduce UKEDR, a deep learning framework that integrates knowledge graph embedding and pre-training strategies to overcome the intractable cold start issue, achieving superior performance and interpretability in drug repurposing.

Kun Li, Jiacai Yi, Qing Ye et al. · 1 citation
#machine learning Open access Jun 2025

HiCLR: Knowledge-Induced Hierarchical Contrastive Learning with Retrosynthesis Prediction Yields a Reaction Foundation Model

HiCLR is the first foundation model that can be broadly applied to various synthesis-related tasks, and it achieves state-of-the-art performance in reaction classification, reaction condition recommendation, reaction yield prediction, synthesis planning, and even molecular property prediction.

Jialu Wu, Yiheng Zhu, Xiaorui Wang et al. · 0 citations
#machine learning Open access Jul 2025

RSGPT: a generative transformer model for retrosynthesis planning pre-trained on ten billion datapoints

RSGPT, a generative model pre-trained on ten billion data points, achieving state-of-the-art performance for synthesis planning, and introduces reinforcement learning to capture the relationships among products, reactants, and templates more accurately.

Yafeng Deng, Xinda Zhao, Hanyu Sun et al. · 18 citations · ⚡2
#machine learning Open access Jun 2025

AntiBMPNN: Structure‐Guided Graph Neural Networks for Precision Antibody Engineering

Antibodies are crucial for medical applications, yet traditional methods for designing sequences are inefficient. This study introduces AntiBMPNN, an advanced deep‐learning framework that leverages an antibody‐specific 3D dataset, a fine‐tuned message‐passing neural network (MPNN), a frequency‐based scoring function, and AlphaFold 3 to achieve highly accurate antibody sequence design. AntiBMPNN surpasses ProteinMPNN with a perplexity of 1.5 and over 80% sequence recovery. Its scoring function, combined with AlphaFold 3, effectively prioritizes sequences based on structural recovery, positional stability, and biochemical or complex properties. Experimental validation highlights a 75% success rate in single‐point antibody design. AntiBMPNN consistently outperforms AbMPNN, AntiFold, and ProteinMPNN in designing complementarity determining regions (CDR) 1‐3, yielding stronger binding affinities. For CDR1 of huJ3 (anti‐HIV nanobody), it achieves a half maximal effective concentration (EC₅₀) of 9.2 nM (nanomolar), better than ProteinMPNN (135.2 nM) and AntiFold (59.3 nM), and comparable to AbMPNN (6.6 nM). For CDR2 of the D6 nanobody (targeting CD16), AntiBMPNN reaches 0.3 nM, outperforming AbMPNN (2.3 nM), AntiFold (0.7 nM), and ProteinMPNN (0.7 nM). In CDR3 of huJ3, it achieves 1.7 nM, surpassing AbMPNN (51.2 nM), with no detectable activity from AntiFold or ProteinMPNN. These findings confirm that AntiBMPNN‐designed sequences for J3 and D6 outperform the originals, highlighting its potential to improve therapeutic antibody design.

Ze-Yu Sun, Jiayi Yuan, Divya Jaiswal et al. · 9 citations
#machine learning Open access Aug 2026

AI-driven PROTAC design overcomes oncogenic resilience by eliminating the CLIP1-LTK fusion protein.

The discovery of CAP-Gly domain-containing linker protein 1(CLIP1)-Leukocyte tyrosine kinase (LTK) as an oncogenic fusion reveals a unique dependency not only on LTK kinase activity but also on CLIP1-mediated multimerization, a noncatalytic function that drives oncogenic signaling. While this fusion is currently targeted with anaplastic lymphoma kinase inhibitors, their exclusive focus on kinase inhibition leaves the scaffolding function intact, necessitating a complete protein clearance strategy. Here, we report the AI-guided development of a first-in-class proteolysis-targeting chimera (PROTAC) designed to selectively degrade the CLIP1-LTK fusion protein. By integrating deep learning models for ternary complex prediction with structure-based molecular optimization, we designed DCL05, an orally bioavailable degrader of CLIP1-LTK fusion protein, achieving picomolar degradation potency (DC50 = 40 pM) and robust antitumor activity. DCL05 consistently outperformed existing kinase inhibitors across a broad spectrum of LTK resistance-associated mutations, both in vitro and in vivo. Collectively, our study explores resistance-associated contexts of LTK and establishes a structure-guided PROTAC development pipeline, providing a promising therapeutic strategy for overcoming acquired resistance in kinase-driven cancers.

Shicheng Chen, Haiting Duan, S. Zhong et al. · 0 citations
#machine learning Open access Apr 2026

Accurate and task-agnostic modeling of enzymatic reactions through multimodal relational learning

ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data, and demonstrates its potential as a versatile and effective tool for enzyme catalysis research.

Yuansheng Huang, Lanqing Li, Wenjia Qian et al. · 2 citations
#machine learning Open access Jul 2026

BBBP-Atlas: Unified Interpretable Modeling of Blood–Brain Barrier Permeability across Small Molecules and Peptides

Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.

Xin Shen, Qun Su, Hao Luo et al. · 0 citations
#machine learning Open access Jun 2026

Targeting the intrinsically disordered AR-NTD through a machine learning-based enhanced sampling workflow

Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.

Kai Zhu, Huating Wang, Jintu Zhang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.