This work introduces mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints and results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics.
Abstract
Motivation With the rapid development of AI methods that predict protein structures from sequence, understanding the structure-function relation increasingly depends on quantitative structural descriptors that are both biologically meaningful and scalable to large datasets. Here, we introduce mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints. Results By employing only three such metrics across all protein structures in the Protein Data Bank, we represent the proteome structural space in a continuous three-dimensional space. Distances within this space capture structural similarity and correlate with functional similarity. We find that the mathematical entanglement based landscape of protein structural space diversifies with the evolutionary expansion of protein function across species. Moreover, this continuous representation reproduces CATH classifications with high accuracy for major structural classes. These results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics. Availability Data used in this study are available in the Protein Data Bank. Details of the machine learning model used can be found in https://github.com/roshitac/CATH_Classification-. Contact Banu.Ozkan@asu.edu, Eleni.Panagiotou@asu.edu Supplementary information Supplementary data are available at Journal Name online.
Support Field Neural Representation Learning (SF-NRL), a topology-guided approach that integrates persistent homology(PH), spatial density estimation, and geometric deep learning to infer residue-wise support directly from protein structures, is introduced.
It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.
R. Khalil, Elena Frasnetti, Han Kurt et al.· Journal of Physical Chemistr...· 0 citations
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
Mohammad Abdulqader, Victor Muñoz· Protein Science· 0 citations
A fundamental question in structural biology centres around understanding protein evolution. Key to this process is mutational robustness, defined as the protein fold’s ability to absorb sequence changes without collapsing its structure. Here, we show that robustness is systematically shaped by simple features such as protein size, geometry, and oligomeric state. We used Foldseek-identified (structural) homologs to quantify family size across monomers and higher homo-oligomers. We found that proteins in larger families are consistently larger in size, more compact in atomic density, and less exposed to solvent. Strikingly, homo-oligomers occupy systematically larger families than monomers, revealing quaternary structure itself as a driver of mutational tolerance, not merely a functional supplement. This signature of robustness can be further linked to increasing functional complexity in proteins; those with adaptive, multifaceted biological roles belong to larger structural families than those with specific roles, thereby linking structural flexibility directly to evolutionary versatility. In short, simple yet overlooked features of protein geometry can explain mutational robustness and evolvability, offering a structural rationale for why certain protein families have diversified extensively while others remain in evolutionary stasis. eTOC blurb (short summary) Why do some protein folds diversify into thousands of variants while others remain rare and rigid? This work shows that the answer lies in geometry: proteins with denser cores, larger size, and higher-order oligomeric assembly tolerate mutations more readily, occupy larger structural families, and support more versatile biological roles. This reveals that protein size, shape, and self-assembly, not just sequence, are fundamental drivers of evolvability.
By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, it is shown that the trained model produces physically structured conformational landscapes, appearing to encode structural constraints that extend beyond what unperturbed inference reveals.
An improved force field is developed, derived from its parent, Amber ff24EXP-GA, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins.