Protein phosphorylation regulates nearly every cellular process, yet most of the hundreds of thousands of human phosphosites remain functionally uncharacterized. Rather than prioritising phosphosites by conservation or structural features, here we use the abundance of an interaction partner as a readout of whether a phosphosite affects that interaction. This idea exploits the fact that subunits of stable complexes are often degraded when unbound. Here, we apply a nested linear regression model to pan-cancer data from 1,006 tumours, while controlling for transcriptional and other covariates. We identified 6,160 associations between 3,038 phosphosites and the abundance of interacting proteins, including several known interaction-regulating sites. Mapping these onto AlphaFold-predicted complexes placed 239 sites at interaction interfaces, while another 402 were linked to compartment-specific localisation, indicating that phosphorylation can also tune interactions by relocating proteins between compartments. Affinity-purification mass spectrometry of NKAP and NUF2 phosphosite mutants experimentally supported some of these predictions. Together, this framework reveals a widespread coupling between phosphorylation and interaction-dependent protein abundance and provides a prioritized, structure-informed resource for characterizing the human phosphoproteome
Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.
It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.
O. Follonier, Yan Liu, Pablo Campomanes et al.· bioRxiv· 1 citation