Short The SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. Long Background The SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. Results Compared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. Conclusions In this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.
V. Vasylieva, Enrico Massignani, Tine Claeys et al.· bioRxiv· 0 citations
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein–protein, protein–ligand, and host–pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voilà applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al.· bioRxiv· 0 citations
Protein phosphorylation is a central regulatory mechanism controlling protein activity, interactions, and cellular signalling, and its dysregulation is implicated in numerous diseases. Advances in mass spectrometry–based phosphoproteomics have led to a rapid expansion in the number of reported phosphorylation sites; however, interpretation of these data remains challenging due to fragmented evidence, limited structural context, and the lack of uniform experimental provenance across resources. Interpretation is further complicated by the fact that the biological meaning of reported phosphosites can vary substantially across tissues, cell lines, perturbations, and disease settings. Here, we present a major update of Scop3P, a proteomics-informed knowledgebase that contextualizes human phosphorylation sites within integrated sequence, structural, biophysical, evolutionary, and mutational frameworks. The current release incorporates uniformly reprocessed human phosphoproteomics data from 116 PRIDE datasets alongside curated UniProt annotations, retaining peptide-spectrum matches, site localization confidence, and direct links to primary mass spectrometry evidence via Universal Spectrum Identifiers. This integration yields 152,350 unique serine, threonine, and tyrosine phosphorylation sites across 16,533 human proteins, supported by full experimental provenance. Beyond site identification, Scop3P provides residue-level contextual annotations derived from experimentally determined protein structures and proteome-wide AlphaFold models, enabling near-complete structural coverage of phosphorylation sites. Structural context is further complemented by residue-level biophysical, evolutionary, and mutational annotations, supporting integrated assessment of phosphorylation in functional and disease-related settings. The current release also introduces residue interaction network representations derived from AlphaFold-predicted structures, capturing spatial connectivity and local interaction environments of phosphorylation and mutation sites. A redesigned web interface enables interactive exploration through coordinated 1D, 2D, 2.5D, and 3D visualizations, peptide-level coverage views, and direct access to original spectra via PRIDE. By bridging experimental phosphoproteomics with structural, functional, and disease-related context, Scop3P provides a scalable and provenance-aware resource for phosphosite interpretation, hypothesis generation, and data-driven modelling of phosphorylation-dependent regulation.
P. Ramasamy, Natalia Tichshenko, Adrián Díaz et al.· bioRxiv· 0 citations