Skip to content
Open access

Reliability of AI Methods in Drug Discovery: Evaluation of Boltz‑2 for Structure and Binding Affinity Prediction

Jul 2026 · Journal of Chemical Theory and Computation · Vol 22, pp. 7811 - 7824 · 1 citation · 53 references
Medicine

TL;DR

An extensive evaluation of Boltz-2 using two large-scale data sets shows that Boltz-2 lacks the energetic resolution required for lead identification, highlighting the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.

Abstract

Despite continuing hype about the role of AI in drug discovery, no “AI-discovered drugs” have so far received regulatory approval. Here we assess one of the latest AI-based tools in this domain. Boltz-2, a recently developed biomolecular foundation model, aims to bridge the gap between AI efficiency and physics-based precision through a joint “cofolding” approach. In this study, we provide an extensive evaluation of Boltz-2 using two large-scale data sets: 16780 compounds for 3CLPro and 21702 compounds for TNKS2. We compare Boltz-2 predicted structures with traditional docking and binding affinities with binding free energies derived from the physics-based ESMACS protocol. Structural analysis reveals significant global RMSD variations, indicating that Boltz-2 predicts multiple protein conformations and ligand binding positions rather than a single converged pose. Energetic evaluations exhibit only weak to moderate correlations across the global data sets. Furthermore, a focused analysis of the top 100 compounds yields no significant correlation between the Boltz-2 predictions and the binding free energies from fine-grained ESMACS, alongside frequently observed saturation-state errors in Boltz-2 predicted ligand structures. Our results show that Boltz-2 lacks the energetic resolution required for lead identification. These findings highlight the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.

Read PDF

Similar papers

Open access Aug 2026

Evaluating molecular docking for binding affinity predictions: a systematic analysis of key parameters and the utility of AlphaFold2 structures for the Schrödinger dataset

Molecular docking is one of the most established methods in computational drug discovery, due to its balance of speed and accuracy. However, the accuracy of docking results depends on a number of different parameters, and systematic reference data for comparisons to more advanced methods for binding affinity prediction are still scarce. This study assesses the impact of key parameters on the accuracy of binding free energy estimates from docking, using nine benchmark systems with 278 high-affinity ligands. Using the Molecular Operating Environment (MOE), we evaluated combinations of three receptor structures (two crystal structures, one AlphaFold2 model), two force fields, two scoring functions, two receptor flexibility settings, and two statistical evaluation schemes. The performance of the docking approaches is measured based on the squared Pearson’s correlation coefficient (R²), the root mean square error (RMSE) with respect to the experimental binding affinities, as well as the mean signed error (MSE) and Kendall’s tau for individual targets and the full dataset. The results show that the scoring function and the protein structure are the most important factors for binding affinity accuracy in rigid docking with the MOE software. Amber10:EHT and MMFF94x force fields had the same average Rmean2 value, but Amber10:EHT had a lower average RMSEmean. AlphaFold2 protein models yielded lower binding affinity accuracy and higher errors compared to experimental crystal structures, although induced fit docking improved results. Using the original benchmark, we also compared several docking programs. DOCK6 and MOE performed best, with mean R² values of about 0.49 and 0.40, respectively. The remaining docking programs did not outperform a molecular weight regression baseline. For a subset of four targets (CDK2, JNK1, P38, TYK2) evaluated in previous work, the performance of the optimized DOCK6 and MOE protocols produced correlation coefficients similar to those reported for certain MM/PBSA, FMO, and Boltz2 implementations evaluated on the same target subset. This raises questions about potential dataset biases, the structural preparation, or the implementation of those methods. Docking therefore should be considered as an important and computationally inexpensive reference baseline for binding affinity prediction.

Konstantinos Tornesakis, J. Essex, Paul A. Cox et al. · 0 citations
Open access Jul 2026

Analysing open-source protein folding models for nanobody binding prediction

These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.

Yannick Vogt, Rebekka Roßberg, Jan Habermann et al. · 1 citation
2026

Open-Sourced In Silico Drug Screening.

This chapter describes a structure-based computational approach to perform high-throughput ligand screens of chemical libraries using open-source software programs and illustrates this workflow with the enzymatic molecular target NAD(P)H:quinone oxidoreductase1 (NQO1), which is overexpressed in a number of human solid tumors.

Audrey G. Fikes, Melissa C. Srougi · 0 citations
Jul 2026

Bridging between Structure-Based and Data-Driven Affinity Prediction.

This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.

Ažbeta Kubincová, David L. Mobley · 1 citation
Open access Jul 2026

Computational Lead Optimization on BACE1: Relative Binding Free Energy Perturbation as the Terminal Refinement Layer

Structure-based drug discovery is known to apply computational methods in a tiered hierarchy, with each layer narrowing the candidate set and refining the binding picture before committing to the next, more expensive step. We present a four-tiered computational benchmarking study evaluating five engines against a panel of 36 compounds targeting β-secretase 1 (BACE1), a validated Alzheimer’s disease target with extensive co-crystal ground truth. This study evaluates Flexible Docking and Boltz2 Cofolding as the primary tier, followed by Ensemble Docking, and then Protein-Ligand MD with MM/PBSA and MM/GBSA post-processing. This is then concluded with Relative Binding Free Energy Perturbation (RevFEP) as the terminal refinement layer. Each method was benchmarked against the experimental binding free energies derived from the co-crystal structures spanning −7.85 to −11.35 kcal/mol. Our findings revealed that Flexible Docking reproduced the co-crystal binding mode for 35 of 36 ligands (97.2% within 2.0 Å RMSD) but did not rank potency at this resolution. Boltz2 CoFolding provided an orthogonal structural cross check with a receptor backbone RMSD of 0.293 Å against the experimental co-crystal structure. Ensemble Docking identified the optimal receptor conformation for downstream FEP setup. MD with MM/GBSA decomposition identified van der Waals complementarity as the primary potency driver (Pearson r = +0.855, R2 = 0.732 on a 10-compound subset). RevFEP delivered the highest affinity correlation of any method (Pearson r = +0.662, R2 = 0.438, Spearman ρ = +0.624, mean absolute error 1.02 kcal/mol across all 36 ligands), resolving potency differences within a narrow 3.5 kcal/mol congeneric window that no other engine could discriminate. We characterize what each engine contributes independently and where RevFEP delivers signals no other engine achieves.

Kristoffer Alejo, Christopher Korban, Christian Chung · 0 citations