Skip to content
Open access

Systematic contextual biases in SegmentNT potentially relevant to other nucleotide transformer models

Jul 2026 · Nucleic Acids Research · Vol 54 · 0 citations · 58 references
Medicine

TL;DR

It is demonstrated that nucleotide position within the input sequence alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency and identify potential approaches to account for these biases.

Abstract

Abstract Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT’s prediction probabilities, revealing an intrinsic bias potentially linked to the model’s training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.

Read PDF

Similar papers

Open access Sep 2026

A token-pruning framework enables efficient representation of the human genome for RNA modification analysis

Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings and structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer.

Wen-Jia Gao, Jun-Lei Yu, Jun-Ru Jin et al. · 0 citations
#small language model Open access Sep 2026

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5′-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanisti...

A. Duggan, M. Newman, David R. McMillen · 0 citations
Book Open access Aug 2026

MotRNA: Encoding RNA Motifs via Explicit N-gram Memory

RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling glo...

Xiang-Yu Ji, Xin Wang, Yang Zhang et al. · 0 citations
Open access Aug 2026

Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT

Motivation RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair...

M. Hossain, MD. Roqunuzzaman Sojib, Md Toki Tahmid et al. · 0 citations
Open access Sep 2026

Towards Sparse Causal Features for Zero-shot Mutation Effect Prediction in a Protein Language Model

Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patch...

Saishradha Mohanty, Manya Phutela, A. G. Green · 0 citations
#artificial intelligence Preprint Sep 2026

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...

Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.