Skip to content
Preprint

PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

This work forms the specialization of pretrained virtual screening models to individual pockets as a test-time adaptation problem and proposes PETA, a parameter-efficient framework that directly adapts pretrained model at test time and outperforms both pretrained and fully retrained baselines while updating only the LayerNorm parameters.

Abstract

Accurately ranking active ligands for a target protein pocket from massive chemical libraries remains a central challenge in virtual screening. DrugCLIP and its recent extensions substantially accelerate this process by encoding protein pockets and molecules into a shared embedding space. Despite this progress, further performance improvements typically require retraining the entire model, incurring substantial computational overhead and making target-specific customization inefficient. In this work, we formulate the specialization of pretrained virtual screening models to individual pockets as a test-time adaptation problem and propose PETA, a parameter-efficient framework that directly adapts pretrained model at test time. Given a target pocket, PETA constructs pocket-specific negatives through molecular diffusion and chemical validity filtering, and further moves them toward the reference ligand retrieved from structural databases via embedding-space mixup to create more challenging ranking tasks. A ranking objective then places greater emphasis on suppressing high-scoring invalid candidates that could contaminate the top-ranked screening results, providing structured supervision for lightweight adaptation. Experiments across diverse benchmarks demonstrate that this lightweight, pocket-specific adaptation outperforms both pretrained and fully retrained baselines while updating only the LayerNorm parameters, which account for approximately $0.03\%$ of the full model.

View source

Similar papers

Review Open access Jul 2026

BoltzMol-1: Towards Reliable Virtual Screening for Fast and Cost-Effective Hit Discovery

This work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles, and introduces a suite of ADMET models for kinetic solubility, lipophilicity, and Caco-2 permeability to improve developability at the point of selection.

Noah Getz, Geoffrey Smith, Avene Colgan et al. · 0 citations
Aug 2026

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Jaeoh Shin, K. Joo, Jejoong Yoo · 0 citations
Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, Matthew B. Soellner, Charles L. Brooks · 0 citations
Open access Jul 2026

Prioritising search for virtual screening via preliminary interpretable low-feature likelihood-based rankings of drug-target activity measures.

BACKGROUND Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance. RESULTS This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class. CONCLUSIONS By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.

Riccardo Curcio, Toni Mancini, Enrico Tronci · 0 citations
Jul 2026

Abstract A024: FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors

Accurate structure-based virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on either lower-fidelity scoring functions or sampling-based strategies, which can limit predictive accuracy and introduce biases in the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of a high-accuracy structure-based model (Boltz-2) into a computationally efficient sequence-based surrogate. By training on ∼1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. We applied this framework to histone deacetylase 11 (HDAC11), the sole class IV member of the histone deacetylase family and epigenetic regulator implicated in tumor progression and therapy resistance, yet remains chemically underexplored with relatively few inhibitors available. Compared with a random background (N = 1.85 million), FastBindRank effectively enriched high-confidence binders, with higher predicted binding probabilities (Cliff’s δ = 0.97) and lower predicted log10(IC50) values (Cliff’s δ = −0.75). Re-scoring and physicochemical filtering yielded 1,262 high-confidence candidates and 528 structurally diverse representatives. Under a comparable computational budget, our approach achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. To interpret the structural patterns underlying model predictions, SHapley Additive exPlanations (SHAP) analysis was performed on the top-ranked candidates (N = 500) which revealed a subset of Morgan fingerprint bits with high contributions, indicating that model predictions are driven by specific local chemical environments The framework’s predictive accuracy was experimentally validated for two novel compounds using an in vitro HDAC11 enzyme activity assay with panobinostat and fimepinostat (both FDA approved pan-HDAC inhibitors) as positive controls. Both novel compounds showed HDAC11 inhibitory activity similar to or higher than the positive controls. The IC50 values were 1.3 µM and 14.8 µM, respectively, which compare favorably with Panobinostat and Fimepinostat that had IC50 values of 20.8 µM and 3.3 µM, respectively. These results provide experimental support for the predictive capability of FastBindRank and demonstrate that large-scale structure-guided prioritization can identify functionally active novel compounds, offering a practical path for candidate discovery from ultra-large virtual screening. Jiawei Dai, Yueyue Wang, Naing Lin Shan, Marco Mariani, Zimeng Yu, Qin Yan, Lalit Golani, Yulia Surovtseva, William Lee, Lajos Pusztai. FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A024.

J. Dai, Yueyue Wang, N. Shan et al. · 0 citations