Structural investigation of re-docked actives showed that re-ranked poses were more native-like, with improved binding-site occupancy, reduced centroid displacement, and greater recovery of co-crystal interactions.
Abstract
False positives in virtual screening often arise when a single docking score or top-ranked pose is treated as sufficient evidence for binding. We extend the previously introduced ProDock software from a database-backed docking platform into a rank-resolved, multi-engine workflow for automated preparation, docking, pose analysis, and optimized re-ranking. The extended workflow combines local docking with GNINA and global docking with DiffDock with pose-level descriptors, namely binding-site occupancy, ligand localization, interaction-fingerprint similarity, and steric clash counts, together with Optuna -based threshold optimization. Across 43 DUDE-Z targets, the archived benchmark outputs reported higher enrichment values for CNN-based GNINA scores after optimization. CNNaffinity PR-AUC changed from 0.197 to 0.294 and LogAUC from 0.708 to 0.763, whereas empirical affinity ROC-AUC changed from 0.770 to 0.758. Structural investigation of re-docked actives showed that re-ranked poses were more native-like, with improved binding-site occupancy, reduced centroid displacement, and greater recovery of co-crystal interactions. The extension provides a reproducible framework for combining complementary docking engines with interpretable pose-level metrics before hit selection, thereby aiding the identification of true-positive candidates in virtual screening.
Overall, this end-to-end workflow efficiently compresses large libraries into a confirmed micromolar CXCR4 hit while lowering computational and infrastructure barriers for early-stage discovery.
Results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds.
Jongkeun Choi· International Journal of Mol...· 0 citations
Covalent virtual screening requires ranking compounds according to both noncovalent recognition and their ability to adopt a reaction-competent geometry with an appropriately reactive warhead. Here, we introduce BCover, a reaction-aware scoring suite that combines pre-reactive docking with quantum-chemistry-derived ligand reactivity descriptors, the electrostatic properties of the protein pocket. Quantum-chemical descriptors are aggregated using a nonlinear tree-based scoring model. BCover was evaluated retrospectively on the COValid benchmark, comprising nine targets and ten reactive sites, and compared with AutoDock, DOCK6, DOCKovalent, AlphaFold3 with Rosetta rescoring, and AlphaFold3 confidence-based ranking. BCover achieved an average adjusted LogAUC of 31% (57% max), an average ROC-AUC of 0.88 (0.96 max), an average EF1 of 18 (33 max). Its average LogAUC exceeded those of the classical docking methods and AlphaFold3-Rosetta, although AlphaFold3-mPAE provided the strongest overall enrichment. At an average runtime of about 10~s per ligand, BCover was approximately 25-fold faster than the evaluated AlphaFold3 workflows and achieved the highest average time-adjusted virtual-screening productivity index. Redocking experiments further showed that the method recovered near-native ligand conformations. These results demonstrate that combining docking-derived geometry with ligand local electronic reactivity and pocket electrostatics provides an efficient and interpretable strategy for covalent ligand prioritization. BCover is intended as a high-throughput screening method that complements more computationally demanding QM/MM and free-energy calculations during subsequent lead optimization.
Molecular docking and co-folding engines are widely used to prioritize compounds for wet-lab validation, yet their accuracy is known to vary substantially across protein targets for reasons that remain only qualitatively understood. Here we benchmark six docking and co-folding engines (RevDock, DiffDock, Boltz2, AutoDock-GPU, rDock, and PandaDock) across 14 protein families, evaluating scoring power, ranking power, docking power, and physical validity. Rather than treating engine performance as protein-family-specific, we classify all 14 families into six mechanistic groups according to which of four scoring-function simplifications, rigid receptor, pairwise additivity, fixed point charges, and implicit solvent, is most severely stressed by that family’s binding site. This framework helps explain, rather than simply describe, where each engine succeeds or fails: RevDock’s CNN rescoring layer mitigates the pairwise additivity and fixed-charge limitations relative to physics-only scoring, achieving the highest overall pose accuracy (73.3% of poses ≤ 2.0 Å RMSD), while Boltz2’s sequence-based co-folding bypasses the rigid-receptor assumption and achieves comparable affinity correlation (mean Pearson r ≈ 0.60 for both engines). PandaDock, run with expanded conformational sampling, matches RevDock on pose accuracy (72.1% of poses ≤ 2.0 Å, lowest median RMSD at 0.96 Å) and exceeds AutoDock-GPU on affinity correlation (mean r = 0.460), indicating that the performance of a physics-based scoring function is limited as much by search adequacy as by the scoring function itself. These results suggest that engine selection for a docking or co-folding campaign should be guided less by an engine’s aggregate benchmark ranking and more by which of these four structural and physical characteristics dominate the target of interest.
Kristoffer Alejo, Sarah Fisher, Tejaswan Kalluri et al.· bioRxiv· 0 citations
“NextTopDocker” is presented, a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank.
Cao-Minh Truong, Pedro J. Ballester, O. Taboureau et al.· Journal of Medicinal Chemist...· 0 citations
A quantitative framework for understanding how docking performance responds to methodological improvements has been lacking is developed by modeling large-scale experiments from three previously published docking campaigns, providing an objective basis for benchmarking and comparing virtual screening approaches.
Laust Moesgaard, Brian K. Shoichet, Olivier Mailhot· Journal of Chemical Informat...· 1 citation