It is suggested that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.
Abstract
Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface-quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-K candidates from already generated binder pools using a shared panel of precomputed proxy scores. On the 10-target held-out split, averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571 Recall@10. On the 3-target held-out subset comprising Nipah, RBX1, and TREM2, target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10 and 0.583 NDCG@10. These results suggest that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.
General-purpose frontier language models are being increasingly utilized for protein-design work, yet their ability to understand and evaluate variant effects remains unclear. Here, we introduce PG-LLM, a benchmark comprising 276 protein-variant prioritization tasks: 217 from ProteinGym and a temporally held-out set of 59 from recently published studies. Each task follows the same format: a language model is asked to rank a list of variant sequences given only the wild-type protein sequence and an assay description with no access to tools, multiple-sequence alignments, or protein structures. We evaluate thirteen language models and 95 published protein predictors on the same variants with the same evaluation metric. Claude Opus 5 (Max) and GPT 5.6 Sol (Max) are the best performing LLMs with Spearman correlations of ρ = 0.406 and 0.402 respectively. Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at ρ = 0.411, but remains below the leading predictor VenusREM at ρ = 0.523. We observe that variant-ranking performance scales with test-time compute across GPT, Claude, and Gemini models, but gains taper before closing the gap to specialist protein predictors. To address contamination risk, we create a held-out evaluation set with 59 DMS assays from 19 studies whose scores first became public after January 2026. On this set, we observe performance and test time compute scaling trends similar to those on the 217 tasks derived from ProteinGym. PG-LLM shows that tool-free language models capture substantial protein-variant signal, outperforming many sequence-based predictors while remaining below the strongest specialized models.
Rohit Arora, L. Chen, Melissa Du et al.· bioRxiv· 0 citations
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, M. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
Identifying proteins that neutralize snake venom toxins is a critical bottleneck in antivenom development, constrained by the scarcity of experimentally resolved toxin-binder structures and the high cost of wet-lab screening. This paper presents a computational screening pipeline that prioritizes toxinbinder candidates for downstream experimental validation, addressing the challenge of candidate ranking when labeled data is limited and supervised models risk overfitting. A graph reranking algorithm, BinderGraph, is proposed: it propagates frozen ESM2 cosine similarity scores across a binder co-occurrence graph constructed from structural training data, without any learned parameters. Evaluated across 10 random toxin-level splits on 61 unique toxins, Frozen ESM-8M with BinderGraph achieves mean Recall@5 of $0.836 \pm 0.077$, outperforming a trained MLP DuaIEncoder (Recall@5 = 0.804 ± 0.099) and a larger ESM2-650M model. These results demonstrate that biologically informed post-processing of pretrained embeddings is more effective than additional parameters or model scale when training data is scarce. Intended as a first-stage filter for experimental followup such as surface plasmon resonance (SPR) and enzyme-linked immunosorbent assays (ELISA), all comparisons are subjected to multi-split statistical significance testing.
This is the first method to expose GNN-derived attributions to an LLM as evidence for property prediction, and achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task.
Junwoo Park, Minyoung Shin, C. Lee et al.· 0 citations
A modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hint with confidence, and a GNN extracts an instance-specific explanatory subgraph via a necessity-based edge-drop intervention.
K. Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras· 1 citation
TTS-Design is proposed, a test-time compute scaling framework that enhances protein sequence design without retraining models or relying on larger training datasets, and can consistently improve sequence recovery and structural reliability across different backbone models, without retraining or increasing model size.
Zizhe Jin, Yi Zheng, Huan Yee Koh et al.· 0 citations