It is demonstrated that nucleotide position within the input sequence alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency and identify potential approaches to account for these biases.
Abstract
Abstract Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT’s prediction probabilities, revealing an intrinsic bias potentially linked to the model’s training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.
Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings and structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer.
Wen-Jia Gao, Jun-Lei Yu, Jun-Ru Jin et al.· Bioinformatics· 0 citations
Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5′-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanisti...
A. Duggan, M. Newman, David R. McMillen· PLoS ONE· 0 citations
RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling glo...
Xiang-Yu Ji, Xin Wang, Yang Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
Motivation RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair...
M. Hossain, MD. Roqunuzzaman Sojib, Md Toki Tahmid et al.· bioRxiv· 0 citations
Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patch...
Saishradha Mohanty, Manya Phutela, A. G. Green· bioRxiv· 0 citations
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...
Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.