It is argued that the field must move beyond sequence-based splits to ensure that AI-driven discovery translates into successful prospective laboratory research, and that the field must move beyond sequence-based splits to ensure that AI-driven discovery translates into successful prospective laboratory research.
Abstract
Accurate prediction of protein-ligand binding affinity is a crucial goal in structure-based drug discovery, with the potential to significantly shorten development timelines. Recently, a new wave of machine learning models based on co-folding, such as Boltz-2 and IsoDDE, has demonstrated performance that matches or exceeds that of gold-standard physics-based methods like Free Energy Perturbation (FEP). This paper provides a critical assessment of these claims, revealing that current benchmarks are heavily influenced by data leakage, and proposes a new benchmark that explicitly controls for data leakage. We demonstrate that splitting by protein-sequence identity is inherently insufficient to prevent data leakage due to “target mirroring,” in which homologous proteins with low overall sequence identity still exhibit highly correlated binding profiles. Our meta-analysis of documents in the ChEMBL 36 database identifies more than 6,000 such assay pairs and finds that leakage persists for sequence-identity thresholds as low as 0.2, well below the values commonly used in benchmarks today. Additionally, we show that a ligand-only baseline model, which lacks protein structural information, achieves surprisingly high performance on the FEP+ 4 and OpenFE benchmarks (r = 0.66 and r = 0.36, respectively). Our results indicate that current benchmarks tend to reward models for memorizing training data and exploiting localized leakage rather than truly learning biophysical principles. To address this issue, we propose the Novelty-Tiered Affinity Benchmark, in which the test data is partitioned into ligand novelty tiers. In the most challenging tier (Tanimoto similarity < 0.35), ligand-only models perform notably worse (r = 0.14), offering a clear baseline for evaluating genuine generalization. We argue that the field must move beyond sequence-based splits to ensure that AI-driven discovery translates into successful prospective laboratory research.
Novo-1, a coarse-grained cofolding framework for binding- affinity prediction, offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, and demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets.
Nikhil Shenoy, David Errington, Emmanuel Bengio et al.· bioRxiv· 0 citations
Boltz is benchmarked using a curated set of ligand-bound human G protein-coupled receptors from families unseen during training, showing that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP+.
Lichirui Zhang, R. Friesner, Edward B. Miller et al.· npj Drug Discovery· 0 citations
DyAb is a pair-wise representation built on top of a pre-trained protein language model that achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across monoclonal antibodies targeting three different antigens.
J. Lin, Jennifer L. Hofmann, Andrew Leaver-Fay et al.· mAbs· 0 citations
Protein structure predictors achieve high single-state accuracy, but it remains unclear whether they can recover functionally relevant conformational ensembles or account for the presence of ligands and/or binding partners. Here, we benchmark AlphaFold3, Boltz-2, Chai-1, and BioEmu on four canonical multi-state proteins (Pf-MATE, LAO, SecA, and β2AR), quantifying state bias and sampling breadth against experimental reference structures. Models frequently default to a dominant state represented in the PDB; small-molecule ligands have weak or inconsistent effects, while large protein partners drive clear conformational switching between states. Multiple sequence alignment (MSA)-based approaches (AF-Cluster and random subsampling) recapitulate similar biases, indicating that this behavior is not unique to newer architectures. These results underscore current limitations for multi-state protein structure prediction and structure-guided ligand discovery. TOC Graphic
Muhui Ye, Yu-Hong Wang, M. Brogi et al.· bioRxiv· 0 citations
These findings provide practical guidance for integrating open-source protein structure prediction models into AI-driven nanobody discovery pipelines while highlighting the need for improved generalization across antigens.
Yannick Vogt, Rebekka Roßberg, Jan Habermann et al.· Frontiers in Bioinformatics· 0 citations
Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.
P. Bayat, Spencer J. Perkins, Sebastian Clancy et al.· bioRxiv· 0 citations