Distillation is established as a practical strategy for scalable, high-fidelity virtual screening of ultra-large chemical libraries by transferring the predictive power of the structure-based model Boltz-2 into an efficient sequence-based surrogate, FastBindRank.
Abstract
Accurate virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on lower-fidelity scoring functions or sampling-based strategies that can limit predictive accuracy and bias the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of the structure-based model Boltz-2 into an efficient sequence-based surrogate. Trained on ∼1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. Applied to histone deacetylase 11 (HDAC11), FastBindRank substantially enriched high-confidence binders relative to the background chemical space. The lightweight model captured structural patterns associated with predicted binding, revealing structural determinants of binding. Under a comparable computational budget, FastBindRank achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. Experimental validation confirmed the activity of two novel compounds. These results establish distillation as a practical strategy for scalable, high-fidelity virtual screening of ultra-large chemical libraries.
Accurate structure-based virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on either lower-fidelity scoring functions or sampling-based strategies, which can limit predictive accuracy and introduce biases in the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of a high-accuracy structure-based model (Boltz-2) into a computationally efficient sequence-based surrogate. By training on ∼1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. We applied this framework to histone deacetylase 11 (HDAC11), the sole class IV member of the histone deacetylase family and epigenetic regulator implicated in tumor progression and therapy resistance, yet remains chemically underexplored with relatively few inhibitors available. Compared with a random background (N = 1.85 million), FastBindRank effectively enriched high-confidence binders, with higher predicted binding probabilities (Cliff’s δ = 0.97) and lower predicted log10(IC50) values (Cliff’s δ = −0.75). Re-scoring and physicochemical filtering yielded 1,262 high-confidence candidates and 528 structurally diverse representatives. Under a comparable computational budget, our approach achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. To interpret the structural patterns underlying model predictions, SHapley Additive exPlanations (SHAP) analysis was performed on the top-ranked candidates (N = 500) which revealed a subset of Morgan fingerprint bits with high contributions, indicating that model predictions are driven by specific local chemical environments The framework’s predictive accuracy was experimentally validated for two novel compounds using an in vitro HDAC11 enzyme activity assay with panobinostat and fimepinostat (both FDA approved pan-HDAC inhibitors) as positive controls. Both novel compounds showed HDAC11 inhibitory activity similar to or higher than the positive controls. The IC50 values were 1.3 µM and 14.8 µM, respectively, which compare favorably with Panobinostat and Fimepinostat that had IC50 values of 20.8 µM and 3.3 µM, respectively. These results provide experimental support for the predictive capability of FastBindRank and demonstrate that large-scale structure-guided prioritization can identify functionally active novel compounds, offering a practical path for candidate discovery from ultra-large virtual screening.
Jiawei Dai, Yueyue Wang, Naing Lin Shan, Marco Mariani, Zimeng Yu, Qin Yan, Lalit Golani, Yulia Surovtseva, William Lee, Lajos Pusztai. FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A024.
J. Dai, Yueyue Wang, N. Shan et al.· Clinical Cancer Research· 0 citations
This work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles, and introduces a suite of ADMET models for kinetic solubility, lipophilicity, and Caco-2 permeability to improve developability at the point of selection.
This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.
Furyal Ahmed, Matthew B. Soellner, Charles L. Brooks· Journal of Chemical Informat...· 0 citations
It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations
This work forms the specialization of pretrained virtual screening models to individual pockets as a test-time adaptation problem and proposes PETA, a parameter-efficient framework that directly adapts pretrained model at test time and outperforms both pretrained and fully retrained baselines while updating only the LayerNorm parameters.
Jia-Qi Lin, Yinghua Yao, Changran Wang et al.· 0 citations