Skip to content
Open access

WASP: a pipeline for functional annotation prediction based on AlphaFold structural models

Jul 2026 · Nature Communications · Vol 17 · 0 citations · 61 references
Medicine

TL;DR

WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Abstract

Protein function annotation is crucial for understanding biological processes and mechanisms. Traditionally, annotations rely on sequence homology, providing valuable insights but often leaving gaps even in well-characterised organisms. With AlphaFold enabling rapid generation of protein structural models, we can now infer function from three-dimensional shape. Here, we present WASP, a pipeline leveraging structural homology to enhance protein annotation prediction at scale, providing a more comprehensive understanding of protein functions across various organisms. WASP relies on network topology for better accuracy and more robust statistical power. We show that WASP achieves superior F1 scores compared to state-of-the-art sequence-based tools when recovering hidden annotations. On 20 industrially relevant organisms, WASP retrieves annotations for 20-30% of previously uncharacterised proteins. We further demonstrate utility in genome-scale metabolic model curation, identifying native candidates for 75-100% of orphan reactions. WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches. WASP predicts protein functions from AlphaFold structures using network-based structural homology, retrieving annotations for 20-30% of uncharacterised proteins and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Read PDF

Similar papers

Open access Jul 2026

ProtPen combines sequence- and structure-based approaches to facilitate protein function predictions on a proteome-wide scale

ProtPen is an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures and is readily extensible to incorporate additional protein function prediction tools.

Diya Mathai, S. Schulze · 0 citations
Open access Aug 2026

HInt: interaction-based homology discovery through accelerated genome-scale AlphaFold screening

Identifying homologous proteins across deep evolutionary distances remains a major challenge because sequence and structural similarity progressively become undetectable over time. Although protein-protein interactions (PPIs) are often constrained by function and evolution, whether conserved interaction interfaces can provide an independent signal for homology detection has remained largely unexplored owing to the computational cost of proteome-scale interaction prediction. Here we introduce HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling. Using HInt, we establish interaction-based similarity as a third axis of homology detection. We show that conserved interaction interfaces reveal homologous relationships that remain inaccessible to conventional sequence- and structure-based approaches. Application of HInt to both prokaryotic and eukaryotic systems, together with experimental validation, uncovered a previously unrecognised VirB5 pilus-tip protein in the F-plasmid type IV secretion system and a previously unannotated F-box-like protein in the Saccharomyces cerevisiae ubiquitin-proteasome system. By enabling practical proteome-scale interaction screening, HInt provides a general framework for uncovering hidden homologues and expands the conceptual landscape of protein homology inference.

Quentin Rouger, P. Paillard, Manon Thomet et al. · 0 citations
Open access Aug 2026

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 Å, 10 Å, and 15 Å radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels—highlighting the benefit of negative training examples—all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Aug 2026

Scop3P-Toolkit: executable structure-aware workflows linking PTMs, peptides, and mutations to protein function

Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein–protein, protein–ligand, and host–pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voilà applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.

Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al. · 0 citations