EvoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features, provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.
Abstract
The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman’s ρ = 0.823) and the full spike protein (ρ = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor–descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data. Author summary Viruses evolve by accumulating mutations along branching lineages, so the changes a virus is likely to acquire next depend on its current sequence. Computational models that predict protein sequences usually ignore this history: they judge whether a sequence looks plausible in general, not whether it is a plausible descendant of a particular ancestor. We asked whether giving a language model explicit information about evolutionary history would help it reproduce how a viral protein actually changes. Using the SARS-CoV-2 spike protein, for which millions of real sequences have been arranged into a detailed evolutionary tree, we trained a model on pairs of ancestor and descendant sequences together with simple measurements taken from the tree, such as how many mutational steps separate the two. Our resulting model, evoPLM-Tree, benefitted substantially more on the information it was given than a model trained on sequences alone, reproduced where mutations occur across spike, and assigned higher probability to substitutions that laboratory experiments show the protein tolerates. As surveillance data are captured for other pathogens, the same approach could help anticipate mutations in newly emerging viruses.
Protein language models (PLMs) score the effects of amino acid replacements as pseudo-probabilities, which are widely utilised to map protein fitness landscapes. However, because their training data relies on natural amino acid sequences, these models conflate protein structural constraints with nucleotide mutation biases and codon accessibility. Using the rapid emergence of the divergent influenza A H3N2 K lineage as a stress test, we investigate how base PLMs (ESM-2 and ESM-C) versus fine-tuned versions of these models capture mutational processes. We systematically implement a parameter sweep to explicitly couple (or decouple) empirical nucleotide mutational supply from PLM-assessed amino acid substitution pseudo-probabilities across evolutionary forecasting tasks. We find that base PLMs implicitly learn generic nucleotide-level mutational constraints, an effect strongly amplified by virus-specific fine-tuning. Incorporating explicit mutational accessibility significantly improves the binary prediction of observed amino acid changes. Conversely, when predicting the final circulating frequency of variants that have already emerged, adding mutational supply degrades performance, confirming that selection dominates post-emergence dynamics. Additionally, we perform amino-acid-level epistatic scanning to investigate protein structural constraints in the context of genetic background. This indicates the improbable antigenic substitution I160K is dependent on co-occurring S144N and N158D mutations in the H3N2 K lineage. Ultimately, current PLM pseudo-probabilities are a composite metric that conflates protein structural fitness with historical biases in mutational supply. Explicitly decoupling these independent evolutionary processes optimises predictive accuracy for real-world pathogen forecasting and isolates pure protein fitness for synthetic design pipelines.
O. MacLean, Kieran D. Lamb, Spyros Lytras et al.· bioRxiv· 0 citations
ABSTRACT The genomic deluge has pushed viral molecular evolution into a site-resolved era. For antigenically evolving viruses such as influenza and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), dense genomic sampling now supports mutation-annotated phylogenies and per-site estimates of mutation and substitution processes. These data highlight strong effects of sequence context, genomic region, RNA structure, and protein-level constraints that are blurred by classic uniform substitution models. In parallel, accurate structure prediction and emerging structure-aware phylogenetic and machine-learning approaches provide practical ways to map mutations onto three-dimensional constraints, identify structurally plausible escape routes, and interpret evolutionary rate variation through solvent exposure, packing, stability, glycosylation, receptor-binding interfaces, and epitope geometry. Finally, antigenic cartography translates some forms of genetic change into an epidemiologically meaningful phenotype—antigenic distance—while predictive modeling increasingly enables sequence-to-antigenicity inference for variants that have not yet been tested experimentally. Here, we outline a practical framework linking sites, structure, and serology for viruses in which antigenic evolution is a major component of immune escape and lineage turnover; highlight why genetic and antigenic “clocks” can diverge; and discuss how integrating genomic surveillance data, phylogenetics, structural analysis, and predictive modeling could support more prospective variant assessment and improved vaccine and therapeutic design.
Sanni Översti, Spyros Lytras, Shusuke Kawakubo et al.· Journal of Virology· 0 citations
Since its introduction in 1968, Influenza A (H3N2) has undergone continuous antigenic evolution, necessitating frequent vaccine updates. To predict antigenicity and characterize antigenic drift without multiple sequence alignments, we present FluEmbed, a computational framework that leverages protein language models. FluEmbed accurately quantified the antigenic impact of viral evolution from RNA sequences, achieving strong predictive performance against hemagglutination inhibition (HI) assay titers (Spearman correlation: ρ = 0.67–0.80). FluEmbed also outperformed sequence-distance baselines (e.g., Hamming and BLOSUM62) and phylogenetic tree-based models that require sequence alignment. Using this model, we conducted in-silico mutagenesis experiments to identify site/amino acid combinations that differentially impacted antigenicity. To systematically investigate how specific mutations influence immune escape, we defined two classes of mutations: ‘constrained’, where only the most likely amino acid changes at historically mutation-prone sites were considered (thereby limiting the mutation space) and ‘unconstrained’, where all possible substitutions were allowed, providing a full exploration of potential antigenic shifts. Constrained mutations often confer limited antigenic changes, whereas unconstrained mutations exhibit greater escape potential, particularly outside the dominant viral lineages. Notably, 3C.2a was the only major lineage in which constrained and unconstrained mutations showed no significant difference (p ≈ 0.95), suggesting ongoing intra-clade competition rather than inter-lineage antigenic replacement. By enabling rapid, alignment-free antigenic prediction directly from sequence data, FluEmbed could complement traditional HI assays in real-time influenza surveillance and inform vaccine strain selection decisions.
A. Forna, Lambodhar Damodaran, C. Gunning et al.· PLoS Computational Biology· 0 citations
Abstract The quantification of genomic conservation has progressed from foundational statistical modeling of evolutionary rates to state-of-the-art deep learning architectures. However, a major resolution gap remains at the zero-rate origin, where standard selection inference tools fail to distinguish between sites that are invariant due to chance (stochastic invariance) or low substitution opportunity and those that are invariant due to extreme purifying selection. We present B-STILL (Bayesian Significance Test of Invariant Low Likelihoods), a hierarchical Bayesian framework designed to resolve the selective landscape of protein-coding genes near the zero-rate limit. By leveraging gene-level rate distributions (prior calibration) and modeling codon-site-specific substitution opportunities (determined by genetic-code degeneracy and nucleotide substitution biases), B-STILL quantifies the statistical significance of observed stasis. We define a rate-based stasis threshold to identify evolutionary stasis anchors (ESAs)—sites where the upper bound on the evolutionary rate is statistically constrained relative to the background rate of the gene due to extreme purifying selection. Validation against clinical and pathogen datasets confirms that ESAs are strong predictors of biological fitness and pathogenicity. Applying B-STILL across viral and mammalian genomes, we identify thousands of significantly clustered ESAs that map to known functional domains and uncharacterized structural motifs. These results establish B-STILL as a scalable, statistically rigorous framework for high-resolution genomic annotation, converting previously uninformative invariant sites into precise markers of extreme evolutionary constraint.
S. K. Kosakovsky Pond, Hannah Verdonk, Steven Weaver et al.· Genome Biology and Evolution· 0 citations
A rapid expansion of influenza A virus (IAV) genome sequencing has transformed global surveillance but has also created major challenges for interpreting the biological significance of viral mutations, particularly amino acid replacements associated with host adaptation. Resources have been created to support mutation annotation and phylogenetic analysis, but there is a need for a tool that integrates experimentally derived phenotypic evidence with evolutionary context in a framework suitable for users without prior training in bioinformatics. Here, we present the Flu Mutation Explorer, an interactive web application that combines large-scale influenza phylogenies with a manually curated database of reported mammalian adaptation mutations, to enable the exploration and interpretation of IAV genetic variation. The underlying database comprises over 1.5 million publicly available IAV sequences and over 1000 mutations associated with mammalian adaptation. The Flu Mutation Explorer enables users to query protein sequences, visualise amino acid distributions across viral lineages, examine host-specific conservation patterns, and identify adaptation mutation with links to supporting literature. We include case studies which demonstrate the platform’s use in assessing amino acid conservation at sites of interest and in rapidly identifying candidate mammalian adaptation mutations during the ongoing H5N1 panzootic. By integrating genomic, phylogenetic, and functional information into an intuitive interface, the Flu Mutation Explorer lowers the barriers to interpreting influenza sequences for specialists and non-specialists alike.
Laura Mojsiejczuk, Derek W. Wright, R. Gifford et al.· bioRxiv· 0 citations
Mutational biases can influence genome composition, but their contribution to protein evolution remains difficult to quantify. Here we utilize a nearly neutral framework that translates nucleotide mutational spectra into expected amino acid substitution patterns and equilibrium amino acid compositions. Using SARS-CoV-2 as a model system, we show that the viral mutational spectrum explains more than 50% of the variation in observed single-nucleotide amino acid substitutions and predicts the overall direction of proteome-wide amino acid composition change during the COVID-19 pandemic. The predictive power of the model varies with selection regime: effectively neutral and weakly deleterious substitutions conform most closely to the mutational expectation, whereas strongly constrained sites and mutational hotspots show larger deviations. This indicates that departures from the nearly neutral baseline provide a quantitative proxy for purifying and positive selection. Extending the analysis across 34 RNA virus species, we find that positive-sense, negative-sense and double-stranded RNA viruses differ systematically in their mutational spectra, and that these differences are associated with predictable shifts in proteome composition. The same relationship is detectable in RNA-dependent RNA polymerase sequences from more than 77,000 viral species. These results indicate that taxon-specific mutational bias contributes persistently to protein evolution across evolutionary scales.
B. Efimenko, Alexander Voronka, Victoriya Skripskaya et al.· bioRxiv· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.