It is shown that model size, training dataset and stochastic elements can bias the predicted p(sequence) away from real fitness, which clarifies the scaling behavior of protein models on fitness prediction and provides practical guidelines for their application and future development.
Protein language models (PLMs) score the effects of amino acid replacements as pseudo-probabilities, which are widely utilised to map protein fitness landscapes. However, because their training data relies on natural amino acid sequences, these models conflate protein structural constraints with nucleotide mutation biases and codon accessibility. Using the rapid emergence of the divergent influenza A H3N2 K lineage as a stress test, we investigate how base PLMs (ESM-2 and ESM-C) versus fine-tuned versions of these models capture mutational processes. We systematically implement a parameter sweep to explicitly couple (or decouple) empirical nucleotide mutational supply from PLM-assessed amino acid substitution pseudo-probabilities across evolutionary forecasting tasks. We find that base PLMs implicitly learn generic nucleotide-level mutational constraints, an effect strongly amplified by virus-specific fine-tuning. Incorporating explicit mutational accessibility significantly improves the binary prediction of observed amino acid changes. Conversely, when predicting the final circulating frequency of variants that have already emerged, adding mutational supply degrades performance, confirming that selection dominates post-emergence dynamics. Additionally, we perform amino-acid-level epistatic scanning to investigate protein structural constraints in the context of genetic background. This indicates the improbable antigenic substitution I160K is dependent on co-occurring S144N and N158D mutations in the H3N2 K lineage. Ultimately, current PLM pseudo-probabilities are a composite metric that conflates protein structural fitness with historical biases in mutational supply. Explicitly decoupling these independent evolutionary processes optimises predictive accuracy for real-world pathogen forecasting and isolates pure protein fitness for synthetic design pipelines.
O. MacLean, Kieran D. Lamb, Spyros Lytras et al.· bioRxiv· 0 citations
Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Yo Akiyama, Zhidian Zhang, Olivia Tang et al.· Cell· 2 citations
It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.
R. Vinod, Samir Char, Ava A. Amini et al.· bioRxiv· 0 citations
MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability, is introduced.
Mingchen Li, Xiaoran Cheng, Fan Jiang et al.· bioRxiv· 0 citations
Deep learning models that predict molecular phenotypes directly from DNA sequence offer a powerful framework for interpreting genomic variation. Recently, AlphaGenome was introduced as a deep sequence-to-function architecture capable of predicting observations that historically required experiments. While the model has shown high accuracy, it was primarily evaluated on human variants scored against a reference genome. Here, we test performance on mouse data, the other species AlphaGenome was trained on although with fivefold fewer features than human (1,128 versus 5,930). We demonstrate that AlphaGenome’s predictive performance varies considerably depending on the functional task. Specifically, predicted quantitative expression effects are directionally weak and compressed roughly 100-fold relative to empirical benchmarks across both reconstructed-haplotype and single-variant regimes. In contrast, canonical splice-site disruptions are recognized with near-identical accuracy in mouse and human (AUC 0.96 versus 0.98), displaying no cross-species divergence in predicted effect magnitude. We developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants. This demonstrates how GenAI innovations that are still under development can safely be harnessed by wrapping a responsible AI layer around the call to intercept flawed results, thereby adhering to international standards, such as the Australian Voluntary AI Safety Standard (VAISS).
Priya Ramarao-Milne, Suyu Ma, L. Sng et al.· bioRxiv· 0 citations