The authors develop a machine learning classifier through the integration of 25,000 proteomics experiments to construct a wiring diagram of human cells, which enables structural modeling of disease-relevant complexes and establishes a highly accurate protein wiring diagram of the cell.
Abstract
Cellular function is driven by the activity of proteins in stable complexes. Protein complex assembly depends on the direct physical association of component proteins. Advances in macromolecular structure prediction with tools like AlphaFold and RoseTTAFold have greatly improved our ability to model these interactions in silico, but an all-by-all analysis of the human proteome’s ~200 M possible pairs remains computationally intractable. A comprehensive cellular map of direct protein interactions will therefore be an invaluable resource to direct screening efforts. Here, we present DirectContacts2, a machine learning model that distinguishes direct from indirect protein interactions using features derived from over 25,000 mass spectrometry experiments. Applied to ~25 million human protein pairs, our model outperforms previous resources in identifying direct physical interactions and enriches for accurate structural models including ~2500 AlphaFold3 models. Our framework enables structural modeling of disease-relevant complexes (e.g. orofacial digital syndrome (OFDS) complex) offering insights into the molecular consequences of pathogenic mutations (OFD1) and broadly, establishes a highly accurate protein wiring diagram of the cell. Knowledge of the physical interactions of proteins provides mechanistic understanding of their function. Here, the authors develop a machine learning classifier through the integration of 25,000 proteomics experiments to construct a wiring diagram of human cells.
Introduction The functions of proteins are primarily governed by coordinated interactions among amino acid residues throughout their three-dimensional structures. Large-scale determination of protein structures has long been made possible by experimental and computational methods; however, studying complex, dynamic, or multimeric systems remains challenging. Protein contact networks (PCNs) offer a graph-based representation of residue-level interactions and enable the application of network analysis techniques to structural data. Nevertheless, many existing tools mainly focus on creating static networks, which limits analytical flexibility. Methods In this study, we introduce Protein Contact Network Explorer (PCNE), a tool for simple construction, visualisation, and analysis of protein contact networks derived from structure data. Results The tool provides flexible residue contact definitions, the exploration of interactive networks, and the extraction of graph-theoretic measures relevant to understanding protein stability, allosteric communication, and functional organisation. Discussion PCNE supports the analysis of key interaction patterns, facilitating both exploratory and hypothesis-driven research in structural biology. The PCNE can be accessed via https://lactdr5rfibhg9m5tmamwg.streamlit.app/.
Akhurath Ganapathy, Sanjana Vijay Krishnan, Arnold Emerson Isaac· Frontiers in Bioinformatics· 0 citations
The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2--Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation.
George A. Rosenberger, Peng Xue, Isabell Bludau et al.· Molecular Systems Biology· 0 citations
Protein-protein interactions mediate a vast range of cellular functions, requiring diverse modes of binding. While recent years have seen major efforts to chart and classify the protein structure universe, we lack comparable methods to assess and cluster that diversity in interface structure at interactome scale. Here, we present Foldseek-Interface, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces. It matches the accuracy of state-of-the-art tools while running up to 230 times faster. Applying it to all biological assemblies in the PDB, we cluster 3.1 million dimers into 77,167 interface clusters and use this resource to characterise interface diversity, evolution, and pathogen mimicry. Application of Foldseek-Interface to resources of predicted protein complex structures rapidly revealed putatively novel interface types worth further experimental interrogation. Foldseek-Interface and the interface cluster resource are freely available as webservers for search (https://search.foldseek.com/interface) and exploration (https://interface.foldseek.com).
J. M. Strom, Sooyoung Cha, R. Kim et al.· bioRxiv· 0 citations