Skip to content
Open access

Extending protein language models to a viral genomic scale using biologically induced sparse attention

Jul 2026 · GigaScience · Vol 15 · 0 citations · 43 references
Medicine

TL;DR

A long-context protein language model is introduced, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors.

Abstract

AbstractAa Background The transformer architecture in deep learning has revolutionized protein sequence analysis. Recent advancements in protein language models have paved the way for significant progress across various domains, including protein function and structure prediction, multiple sequence alignments, and mutation effect prediction. A protein language model is commonly trained on individual proteins, ignoring the interdependencies between sequences within a genome. However, biological understanding reveals that protein–protein interactions span entire genomic regions, underscoring the limitations of focusing solely on individual proteins.Ab Findings To address these limitations, we propose a novel approach that extends the context size of transformer models across the entire viral genome. By training on large genomic fragments, our method captures putative long-range dependencies consistent with inter-protein relationships and encodes protein sequences with integrated information from distant proteins within the same genome, offering benefits across downstream tasks. Viruses, with their densely packed genomes, minimal intergenic regions, and protein annotation challenges, are ideal candidates for genome-wide learning. We introduce a long-context protein language model, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors. Our semi-supervised approach supports long sequences of up to 61,000 amino acids (aa).Ac Conclusion Our evaluations show improved prediction of masked aa and improved downstream discrimination relative to single-protein models and long-context baselines, with additional validation that our inferred links correlate with independently curated interaction resources.

Read PDF

Similar papers

Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Review Open access Aug 2026

A new dimension in protein-RNA interface prediction: Integrating protein language models and geometric deep learning.

These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.

Rozeena Arif, Alfredo Castello · 0 citations
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 2 citations
Open access Aug 2026

Interpreting Protein Language Models: high attention sites predict functional regions

The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.

Sophia J. Pribus, Russ B. Altman, Gowri Nayar · 0 citations
Preprint Aug 2026

Interpreting Latent Protein Language Model Features with Geometric Annotations

Protein language models (pLMs) encode information about protein sequences which enable downstream tasks such as structure prediction, but their internal representations are not well understood. Sparse autoencoders (SAEs) provide a promising tool to disentangle latent pLM representations into interpretable features, but existing annotation pipelines largely rely on protein-level annotations derived from database labels and LLM annotations of top activating sequences. Such annotations can overlook the localized residue-level and geometric patterns encoded by sparse features. We introduce an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_\alpha$ backbone. Across ESM-2 8M layers, an FDR-controlled discovery analysis shows that local geometry is significantly associated with many SAE features, with varying levels of predictive strength, expanding coverage beyond database and sequence-based methods. In particular, geometry can distinguish SAE features sharing the same database annotation, revealing substructure within known biological labels. A significant portion of SAE features activate on unannotated metagenomic protein sequences enabling us to use our SAE annotations to better understand these sequences. In addition, ablation experiments at the level of contact prediction show that removing found geometric features shifts ESM-2's predicted contact maps in the direction of the descriptor. This provides a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interpretability and structural biology.

S. Setlur, Djordje Mihajlovic, Darrick Lee · 0 citations
Aug 2026

Examine neural network models for protein sequencing based features prediction

The autoencoder framework encodes protein sequence information related to domains, families, and patterns—into a lengthy, sparse binary vector that outperforms other neural network models, including convolutional neural networks, recurrent neural networks, long short-term memory networks, and bidirectional long short-term memory networks.

Biswajit Senapati, Ranjita Das · 0 citations