Skip to content
Preprint

DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction

Aug 2026 · 0 citations
Biology Computer Science

TL;DR

DegradeQuery, a context-aware prediction framework that converts label-missing records into a pretraining signal, is introduced and demonstrates that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.

Abstract

Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.

View source

Similar papers

#machine learning Preprint Sep 2026

ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases

Proteolysis-targeting chimeras (PROTACs) have emerged as a transformative therapeutic strategy that selectively degrades historically''undruggable''targets via the ubiquitin-proteasome system. Despite growing efforts to develop computational predictors of PROTAC degradation activity, existing supervised approaches rema...

Yuansheng Liu, Yu-Fei Ye, Tao Tang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...

Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al. · 0 citations
Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations
Open access Sep 2026

Integrating structural and biological evidence to rerank ESMFold2 protein-protein interactions

Large-scale protein structure prediction enables proteome-wide protein–protein interaction (PPI) screening, but distinguishing biologically meaningful interactions from spurious interfaces remains challenging. Here, we develop a scalable framework combining fast, MSA-free ESMFold2 prediction with PAE-guided domain pars...

Jin-Dou Xie, Ming Li, Yongping Chai et al. · 0 citations
#machine learning Preprint Sep 2026

Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets

Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with...

J. Bernett, A. Spannagl, Joel Ås et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.