Discovery of non-canonical proteins through modification-aware proteogenomics
Short The SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. Long Background The SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. Results Compared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. Conclusions In this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.