Skip to content
Open access

SurroDock: A Deep Learning Surrogate for Accelerated Pre-Docking Ligand Prioritization in Structure-Based Virtual Screening

Jul 2026 · International Journal of Molecular Sciences · Vol 27 · 0 citations · 40 references
Medicine

TL;DR

Results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds.

Abstract

The rapid expansion of make-on-demand and public chemical libraries has made exhaustive docking-based structure-based virtual screening increasingly difficult. This study introduces SurroDock, a lightweight deep-learning surrogate designed to approximate AutoDock Vina docking scores from low-cost two-dimensional molecular features, serving as a practical pre-filter for docking. SurroDock was evaluated for estrogen receptor alpha using two distinct conformations: an agonist-bound (PDB ID: 1GWR) and an antagonist/SERM-bound (PDB ID: 3ERT). The dataset comprised approximately 334,000 unique compounds curated from the NCI Open Database, PubChem, and BindingDB, all docked using a standardized AutoDock Vina workflow. The model was trained on concatenated 2D molecular representations comprising Morgan fingerprints, MACCS keys, RDKit physicochemical descriptors, Vina-inspired ligand descriptors, atom-pair fingerprints, and 2D pharmacophore fingerprints. The docking-score distributions differed substantially between receptor states, with 3ERT exhibiting more favorable scores than 1GWR and weak inter-state score correlation supporting state-specific modeling. Using the integrated Unified-200k training set (200,000 compounds randomly sampled per receptor from the three docked sources), SurroDock achieved strong held-out validation performance, with R2 values of approximately 0.88 for 1GWR and 0.93 for 3ERT. In retrospective screening-style evaluation, SurroDock recovered substantial fractions of Vina’s top-ranked compounds at the top-1% recall (Recall@1%) of approximately 0.57 and 0.61 for 1GWR and 3ERT, respectively, yielding corresponding enrichment factors (EF@1%) of approximately 57-fold and 61-fold relative to random selection. Overall, the results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds. Because SurroDock emulates a docking scoring function rather than experimental binding affinity, its predictions should be used as prioritization aids and complemented by confirmatory docking, pose inspection, and experimental validation.

Read PDF

Similar papers

Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Review Open access Jul 2026

Topological deep learning for drug–target interaction, virtual screening, and docking scoring: a practical, benchmark-driven review

A decision-oriented taxonomy and a benchmark-driven evaluation playbook that specifies minimum standards for splits, metrics, baselines, and ablations to isolate the topological contribution are presented.

Beatriz Suay-García, Antonio Falcó · 0 citations
Aug 2026

NextTopDocker: A Large-Scale Docking-Power Benchmark Reveals Limitations of Current End-to-End Machine-Learning Docking and the Strength of Hybrid Rescoring

“NextTopDocker” is presented, a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank.

Cao-Minh Truong, Pedro J. Ballester, O. Taboureau et al. · 0 citations
Aug 2026

Free energy perturbation and machine learning-assisted identification of potential MAP3K8 hit molecules: a comprehensive structure- and ligand-based studies.

The convergence of docking, dynamics, and free-energy results prioritized PM2, PM3, and PM4 as promising MAP3K8 hit candidates, which require further experimental validation and lead optimization.

M. Islam, A. Iqbal, M. A. Ali et al. · 0 citations