Skip to content
Review Open access

BoltzMol-1: Towards Reliable Virtual Screening for Fast and Cost-Effective Hit Discovery

Jul 2026 · bioRxiv · 0 citations
Biology Medicine

TL;DR

This work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles, and introduces a suite of ADMET models for kinetic solubility, lipophilicity, and Caco-2 permeability to improve developability at the point of selection.

Abstract

We present BoltzMol-1, a small-molecule hit discovery pipeline, centered on an optimized version of Boltz-2, explicitly adapted for prospective discovery. Reliable hit discovery that generalizes across target classes (rather than only the well-characterized families that dominate existing ligand data) would broaden the range of biology accessible to small-molecule intervention and reduce reliance on resource-intensive high-throughput screening. Towards this goal, the system prioritizes compounds for rapid experimental validation by coupling model-driven ranking with streamlined procurement from commercial catalogs. To improve developability at the point of selection, we introduce a suite of ADMET models for kinetic solubility (logS), lipophilicity (logD), and Caco-2 permeability. These models act as an early triage layer, systematically filtering out compounds with unfavorable physicochemical and absorption properties prior to synthesis or purchase. Across a panel of ten targets (most with no representation in the underlying affinity training data) we observe strong prospective performance on challenging systems. Functional actives or binders were identified for 6 of 10 targets, despite modest experimental budgets of 28-96 compounds per target. These results include successes on receptors and enzymes traditionally considered difficult for structure- or ligand-based approaches. Collectively, this work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles. Figure 1: Overview of the prospective virtual-screening campaigns across all targets. For each target, the panel shows the predicted protein-ligand complex together with the number of compounds tested, the number of confirmed actives/binders, and the assays used for screening and follow-up.

Read PDF

Similar papers

Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, Matthew B. Soellner, Charles L. Brooks · 0 citations
Open access Aug 2026

Nesso-1: Accelerating Open-Source Binding Affinity Predictions

Novo-1, a coarse-grained cofolding framework for binding- affinity prediction, offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, and demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets.

Nikhil Shenoy, David Errington, Emmanuel Bengio et al. · 0 citations
Open access Jul 2026

Automated Parallel Synthesis Accelerates Virtual Screening Hit Discovery

Virtual screening (VS) is a powerful approach to exploring a vast chemical space, encompassing libraries of millions to billions of compounds. However, the low hit rates of VS require testing numerous candidates to validate true binders, followed by iterative optimization cycles, which makes experimental validation costly and time-consuming. Here, we report COMBINAUT, an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement. Using a faculty-wide collection of in-house building blocks, the system enables enumeration of over 22.9 million compounds, each designed for parallelized synthesis within 32 h using repurposed solid-phase peptide synthesis equipment. Using this platform, we performed large-scale VS targeting the allosteric pocket of the immuno-oncology target, C–C chemokine receptor 2 (CCR2). Our approach facilitated the rapid synthesis and testing of 100 VS hits spanning diverse molecular architectures. In radioligand binding assays, we successfully validated nine hits with distinct scaffolds, including completely novel CCR2 ligand chemotypes. Iterative hit-to-lead optimization using the automated workflow produced cell-active CCR2 antagonists. This work demonstrates the synergy of automated synthesis and VS, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.

Sean M McKenna, M. Šícho, Cas van der Horst et al. · 0 citations
Jul 2026

Abstract A024: FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors

Accurate structure-based virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on either lower-fidelity scoring functions or sampling-based strategies, which can limit predictive accuracy and introduce biases in the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of a high-accuracy structure-based model (Boltz-2) into a computationally efficient sequence-based surrogate. By training on ∼1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. We applied this framework to histone deacetylase 11 (HDAC11), the sole class IV member of the histone deacetylase family and epigenetic regulator implicated in tumor progression and therapy resistance, yet remains chemically underexplored with relatively few inhibitors available. Compared with a random background (N = 1.85 million), FastBindRank effectively enriched high-confidence binders, with higher predicted binding probabilities (Cliff’s δ = 0.97) and lower predicted log10(IC50) values (Cliff’s δ = −0.75). Re-scoring and physicochemical filtering yielded 1,262 high-confidence candidates and 528 structurally diverse representatives. Under a comparable computational budget, our approach achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. To interpret the structural patterns underlying model predictions, SHapley Additive exPlanations (SHAP) analysis was performed on the top-ranked candidates (N = 500) which revealed a subset of Morgan fingerprint bits with high contributions, indicating that model predictions are driven by specific local chemical environments The framework’s predictive accuracy was experimentally validated for two novel compounds using an in vitro HDAC11 enzyme activity assay with panobinostat and fimepinostat (both FDA approved pan-HDAC inhibitors) as positive controls. Both novel compounds showed HDAC11 inhibitory activity similar to or higher than the positive controls. The IC50 values were 1.3 µM and 14.8 µM, respectively, which compare favorably with Panobinostat and Fimepinostat that had IC50 values of 20.8 µM and 3.3 µM, respectively. These results provide experimental support for the predictive capability of FastBindRank and demonstrate that large-scale structure-guided prioritization can identify functionally active novel compounds, offering a practical path for candidate discovery from ultra-large virtual screening. Jiawei Dai, Yueyue Wang, Naing Lin Shan, Marco Mariani, Zimeng Yu, Qin Yan, Lalit Golani, Yulia Surovtseva, William Lee, Lajos Pusztai. FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A024.

J. Dai, Yueyue Wang, N. Shan et al. · 0 citations
Open access Jul 2026

Prioritising search for virtual screening via preliminary interpretable low-feature likelihood-based rankings of drug-target activity measures.

BACKGROUND Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance. RESULTS This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class. CONCLUSIONS By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.

Riccardo Curcio, Toni Mancini, Enrico Tronci · 0 citations