Skip to content

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.

Jul 2026 · Journal of Chemical Information and Modeling · 0 citations · 29 references
Medicine

TL;DR

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Abstract

The rapid expansion of compound libraries has significantly advanced drug discovery, especially through ultralarge library screenings that provide access to vast chemical spaces. However, the sheer scale of these libraries introduces substantial challenges in the early stages of drug discovery. While it is true that searching larger libraries can improve the hit rate of a virtual screening campaign, as the available chemical space has increased to over 64 billion molecules, identifying relevant information becomes a complex and time-consuming task. Machine learning (ML)-based docking methods and scoring functions offer a potential solution by providing increased speed and scalability. However, much of their reported success is based on benchmark data sets that rely on constructed decoys and ligand sets, which often have hidden biases that can artificially inflate performance and fail to capture the challenges of real-world screening. In this work, we systematically evaluate the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data. By using HTS data sets as a more realistic benchmark, we aim to provide a clearer picture of how ML models perform in practical virtual screening scenarios compared to traditional physics-based docking approaches. Our work highlights the strengths and limitations of current ML methods and offers insights into their reliability and applicability to prospective applications in drug discovery.

View source

Similar papers

Aug 2026

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Jaeoh Shin, K. Joo, Jejoong Yoo · 0 citations
Review Open access Jul 2026

BoltzMol-1: Towards Reliable Virtual Screening for Fast and Cost-Effective Hit Discovery

This work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles, and introduces a suite of ADMET models for kinetic solubility, lipophilicity, and Caco-2 permeability to improve developability at the point of selection.

Noah Getz, Geoffrey Smith, Avene Colgan et al. · 0 citations
Open access Aug 2026

Beyond the Score: Fixed-Budget Benchmarking of Virtual Screening Integration Strategies for Decision-Centric Drug Discovery

Virtual screening (VS) workflows often combine structure- and ligand-based methods; however, their value depends on the number of compounds that can be tested. We benchmarked 20 fixed-budget strategies derived from molecular docking (GNINA CNN score), maximum common substructure (MCS) similarity, and a calibrated machine-learning (ML)-QSAR classifier across five pharmacologically diverse targets. Individual methods, best-rank and worst-rank fusion, mean-rank consensus, and sequential funnels were evaluated at 1%, 5%, and 10% library fractions, with every strategy selecting the same number of compounds. ML-QSAR was the strongest standalone method, recovering 47.6%, 81.6%, and 84.4% of actives at the three cutoffs. At the 1% budget, ML-QSAR achieved the highest mean hit recovery (47.6% recall; 99.2% precision). At 5% and 10%, best-rank fusion of QSAR and MCS produced the highest mean recall (83.2% and 86.4%). Among the sequential workflows, QSAR → MCS achieved the highest 1% hit recovery (45.2 ± 3.3% recall), whereas docking-first funnels consistently underperformed under the default, non-optimized conditions evaluated in this study. Target-level results showed substantial variability in MCS-containing workflows and limited benefits from adding docking without target-specific optimization. Under matched assay budgets, a validated ligand-based predictor or a simple two-method rank-fusion scheme provided the highest observed mean hit recovery without requiring elaborate integration.

Elisabetta Grazia Tomarchio, Rocco Buccheri, A. Rescifina · 0 citations
Open access Jul 2026

Prioritising search for virtual screening via preliminary interpretable low-feature likelihood-based rankings of drug-target activity measures.

BACKGROUND Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance. RESULTS This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class. CONCLUSIONS By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.

Riccardo Curcio, Toni Mancini, Enrico Tronci · 0 citations