PCIPG 2.0 is presented, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity.
Abstract
Abstract Motivation Protein complexes execute cellular functions, yet identifying them from protein–protein interaction (PPI) networks remains challenging because interactomes are incomplete and purely topology-driven clustering often lacks mechanistic interpretability. Here we present PCIPG 2.0, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity. PCIPG 2.0 first fuses multiple omics views to prioritize high-confidence candidate protein associations and enhance the observed interactome, and then learns structure-aware node representations by aggregating residue embeddings on residue graphs, coupled with a PPI-level graph autoencoder to infer latent complex memberships. Results Across five yeast benchmarks, PCIPG 2.0 consistently improves complex recovery over representative baselines and yields predicted complexes with significantly higher Gene Ontology semantic coherence than size-matched random sets. Literature-supported case studies and AlphaFold3-based assembly analyses further suggest that representative predictions are consistent with functionally coherent and structurally plausible protein assemblies. Together, these results suggest that combining multi-omics-driven interactome completion with residue-informed representation learning provides a useful and mechanistically informed framework for protein complex identification under incomplete interactome measurements. Availability The source code is available at GitHub: https://github.com/hyx-1/PCIPG2.0. The complete reproducibility package, including the code, processed data, configuration files and materials required to reproduce the experiments reported in this manuscript, has been archived on Zenodo with an archival DOI: https://doi.org/10.5281/zenodo.20133228. The GitHub repository also provides the README-based usage instructions.
Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.
Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al.· bioRxiv· 0 citations
X-PAIR is presented, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues, and links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.
A pipeline reformulating kinase-substrate modeling as a Bayesian inference problem is presented and it is revealed that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
Protein–protein interaction (PPI) information is distributed across resources that differ in organism coverage, identifier systems, evidence models, confidence scores and access mechanisms, so assembling and comparing evidence for a protein requires source-specific queries, identifier conversion and extensive post-processing. We present KlinkPPI, a web server that retrieves, compares and exports PPI evidence from STRING, BioGRID, IntAct, CORUM, HuRI and Predictomes from a single query. KlinkPPI accepts UniProtKB accessions, NCBI Gene and Ensembl identifiers and gene names, and performs taxonomy-aware mapping to a common identifier space. Users can query individual proteins across all resources available for an organism, or retrieve organism-wide interaction sets. Results are presented per source so that database-specific evidence, annotations and confidence values are retained, while an integrated view exposes coverage and agreement between resources. KlinkPPI deliberately does not merge heterogeneous confidence scores, nor collapse functional associations, complex co-membership, binary interactions and structural predictions into a single consensus network. Results can be exported in PSI-MI TAB 2.8-compatible or Apache Parquet format with user-selected evidence fields. KlinkPPI is freely available at https://rappsilberlab.org/KlinkPPI/ and the source code at https://github.com/Rappsilber-Laboratory/KlinkPPI. Graphical abstract
A. Lutfi, Sukrit Dang, R. Warneke et al.· bioRxiv· 0 citations
The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2--Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation.
George A. Rosenberger, Peng Xue, Isabell Bludau et al.· Molecular Systems Biology· 0 citations
This work presents HGRL-PPIS, a novel hierarchical graph representation learning approach for predicting protein-protein interaction sites that achieves superior performance over competing methods on multiple benchmark datasets, enabling more reliable detection of protein-protein binding residues.