Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.
Abstract
Motivation Continuing advances in genome and metagenome sequencing expand the number of identified conserved protein families that remain functionally uncharacterized and contain domains of unknown function (DUFs). Functional-association resources such as STRING provide biological context, but mostly do not distinguish indirect association from physical interaction. We assessed whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins. Results We generated four structural-prediction cohorts from STRING associations involving DUF-containing proteins and evaluated the predicted complexes using interface ipSAE, average pLDDT and buried surface area. An L2-regularized logistic regression model was trained on an initial cohort of predictions from high-confidence STRING associations to prioritize DUF-containing candidates likely to produce structurally confident AlphaFold 3 complexes. The model was then applied across all 12,535 organisms represented in STRING v12.0, followed by grouping into DUF-family and partner-architecture modules, covering 2,076 unique DUF families. The final L2-model screen contained 12,298 successfully modelled protein pairs, including 1,208 (9.82%) complexes meeting a strict-confidence criterion and 2,433 (19.78%) meeting a more liberal confidence criterion. Two examples suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation. Availability and implementation Predicted structures and associated metadata are available through Zenodo at https://doi.org/10.5281/zenodo.21875362. The model implementation and code used to generate the analyses and figures are available at https://github.com/linoriep/Proteome-scale-structure-prediction-of-DUF-containing-protein-protein-interactions.
Identifying homologous proteins across deep evolutionary distances remains a major challenge because sequence and structural similarity progressively become undetectable over time. Although protein-protein interactions (PPIs) are often constrained by function and evolution, whether conserved interaction interfaces can provide an independent signal for homology detection has remained largely unexplored owing to the computational cost of proteome-scale interaction prediction. Here we introduce HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling. Using HInt, we establish interaction-based similarity as a third axis of homology detection. We show that conserved interaction interfaces reveal homologous relationships that remain inaccessible to conventional sequence- and structure-based approaches. Application of HInt to both prokaryotic and eukaryotic systems, together with experimental validation, uncovered a previously unrecognised VirB5 pilus-tip protein in the F-plasmid type IV secretion system and a previously unannotated F-box-like protein in the Saccharomyces cerevisiae ubiquitin-proteasome system. By enabling practical proteome-scale interaction screening, HInt provides a general framework for uncovering hidden homologues and expands the conceptual landscape of protein homology inference.
Quentin Rouger, P. Paillard, Manon Thomet et al.· bioRxiv· 0 citations
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations
The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Lars A. Eicholt, Lasse Middendorf· bioRxiv· 0 citations
Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Yo Akiyama, Zhidian Zhang, Olivia Tang et al.· Cell· 2 citations
PCIPG 2.0 is presented, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity.
Yixiang Huang, Jiudong Wang, Lei Yang et al.· Bioinformatics· 0 citations
WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.