Abstract Standard probabilistic models of coding sequence evolution effectively identify where and when selection acts but remain agnostic to the mechanistic realization of these forces. We introduce PRIME (PRoperty Informed Models of Evolution), a framework of codon-level maximum likelihood methods—including global (G-PRIME), episodic (E-PRIME), and site-specific (S-PRIME) implementations—that explicitly model amino acid exchangeability as a function of physicochemical properties. By parameterizing attributes such as molecular volume, hydropathy, and secondary structure propensities, PRIME aims to resolve the biophysical basis of selective constraint across both the sequence and the phylogeny. At the site level, S-PRIME leverages an explicit biophysical taxonomy to categorize residues as conserved, neutral, or changing for specific properties, resolving selective signals that are missed by traditional rate-based metrics. Our analysis of a benchmark of 24 diverse datasets and a genome-wide screen of 18,944 mammalian genes demonstrates that consideration of biophysical realism can yield substantial improvements in model fit, acting synergistically with rate variation to explain complex evolutionary patterns. We find that physicochemical constraints at individual sites can be reliably detected in datasets with sufficient information redundancy (substitutions per unique amino acid; AUC=0.91), with sensitivity exceeding 90% in data-rich alignments. E-PRIME reveals a distinct hierarchy in biophysical constraints: while core packing and beta-sheet scaffolds are rigidly conserved, alpha-helix propensity and surface electrostatics serve as the primary substrates for adaptive tuning. Furthermore, PRIME importance weights align with aspects of the primary semantic axes of deep learning representations (ESM-2) and capture key features of experimental fitness landscapes. By transforming abstract evolutionary rates into interpretable biophysical rules, PRIME provides a useful framework for characterizing the mechanistic drivers of protein diversity.
Hannah Kim, Konrad Scheffler, Anton Nekrutenko et al.· Molecular biology and evolut...· 0 citations
Abstract The quantification of genomic conservation has progressed from foundational statistical modeling of evolutionary rates to state-of-the-art deep learning architectures. However, a major resolution gap remains at the zero-rate origin, where standard selection inference tools fail to distinguish between sites that are invariant due to chance (stochastic invariance) or low substitution opportunity and those that are invariant due to extreme purifying selection. We present B-STILL (Bayesian Significance Test of Invariant Low Likelihoods), a hierarchical Bayesian framework designed to resolve the selective landscape of protein-coding genes near the zero-rate limit. By leveraging gene-level rate distributions (prior calibration) and modeling codon-site-specific substitution opportunities (determined by genetic-code degeneracy and nucleotide substitution biases), B-STILL quantifies the statistical significance of observed stasis. We define a rate-based stasis threshold to identify evolutionary stasis anchors (ESAs)—sites where the upper bound on the evolutionary rate is statistically constrained relative to the background rate of the gene due to extreme purifying selection. Validation against clinical and pathogen datasets confirms that ESAs are strong predictors of biological fitness and pathogenicity. Applying B-STILL across viral and mammalian genomes, we identify thousands of significantly clustered ESAs that map to known functional domains and uncharacterized structural motifs. These results establish B-STILL as a scalable, statistically rigorous framework for high-resolution genomic annotation, converting previously uninformative invariant sites into precise markers of extreme evolutionary constraint.
S. K. Kosakovsky Pond, Hannah Verdonk, Steven Weaver et al.· Genome Biology and Evolution· 0 citations
A collection of over 3000 independent, full-length SARS-CoV-2 sequences deriving from posited or confirmed chronic infections is assembled and 14 distinct mutation patterns (MPs) that repeatedly appear in these sequences are described, including four CD8 T cell-escape MPs and two MPs that represent adaptation to tissue compartments outside the upper-respiratory tract.
R. Hisner, Ravindra K. Gupta, Darren P. Martin· bioRxiv· 0 citations