Skip to content
Open access

BoltzOmics: Predicting genetic variant effects on drug binding with Boltz-2

Jul 2026 · iScience · Vol 29, pp. 116797 · 1 citation · 124 references
Medicine

TL;DR

BoltzOmics is an interactive, open-source platform that integrates Boltz-2, a deep learning model for biomolecular structure prediction, to rapidly assess mutation effects on drug binding, and establishes a practical AI-driven framework for accelerating computational drug discovery and advancing precision medicine research.

Abstract

Summary A mechanistic understanding of how genetic variants alter drug-receptor binding is central to precision medicine, drug response prediction, and drug development. Yet, experimental mutation-drug profiling remains slow and expensive, while existing computational approaches often trade accuracy for scalability. We developed BoltzOmics, an interactive, open-source platform that integrates Boltz-2, a deep learning model for biomolecular structure prediction, to rapidly assess mutation effects on drug binding. Starting from amino acid sequences, the workflow queries databases for genetic variants, generates wild-type and mutant protein structures, and screens multiple drugs across variants to predict binding affinity changes. We evaluated BoltzOmics across four targets: hERG, NaV1.5, HER2, and CYP3A4. Predictions achieved Pearson correlations with experimental drug IC50 data up to 0.76 for wild-type proteins and 0.60 for mutants. By enabling scalable, high-throughput assessment of drug-variant interactions, BoltzOmics establishes a practical AI-driven framework for accelerating computational drug discovery and advancing precision medicine research.

Read PDF

Similar papers

Open access Aug 2026

Assessing Computational Models for Pharmacogenomic Variant Interpretation

Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model–based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.

F. Pucci, Pauline Hermans, Matsvei Tsishyn et al. · 0 citations
Preprint Jul 2026

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.

Ali Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami et al. · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 0 citations
Open access Jul 2026

Deep Contrastive Learning for High‐Throughput Prediction of Drug Resistance Mutations from Sequences

ABSTRACT Mutation‐induced drug resistance challenges both pandemic surveillance and drug discovery. While experimental assays are resource‐intensive, current computational predictions remain limited by the scarcity of 3D mutant protein structures. We present DeepMutDTA, a structure‐independent model pre‐trained on 1.5 million data points to predict drug‐target affinity and uncover underlying interaction mechanisms. However, like other sequence‐based approaches, it often falls short in predicting mutant affinities due to the overwhelming sequence similarity between wild‐type (WT) and mutant (MT) targets. To bridge this gap, we introduce SimSiam‐MuTF, a novel fine‐tuning framework to enhance the detection of resistance variants by explicitly aligning latent embedding distances with the corresponding shifts in binding affinity between WT and MT targets. Compared to representative baselines, our model exhibits remarkable robustness across varied sequence identities and unseen data splits, yielding average performance gains of 2.47% (PCC) and 5.10% (SCC) in regression tasks, alongside 4.00% (AUC) and 4.17% (AUPR) in classification tasks. Applications to SARS‐CoV‐2, HIV‐1, and cancer‐related targets highlight its generalization potential and utility in informing therapeutic strategies against drug resistance. Collectively, this robust computational pipeline and fine‐tuning framework deepen our understanding of mutation‐induced resistance and may serve as a powerful platform to accelerate drug discovery against mutant targets.

Xiaowen Hu, Pan Zhang, Shangqian Wu et al. · 0 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Open access Jul 2026

Essentiality-driven prediction of anticancer drug responses in preclinical and clinical contexts

DrGee is presented, an essentiality-centered platform that infers drug sensitivity solely from gene expression profiles, and the built-in DeepEEAA model integrates gene expression, gene essentiality, drug-protein affinity, and drug-gene associations to quantitatively predict IC50 values.

Hongtu Cui, Xiaohui Du, Hai-Xia Guo et al. · 0 citations