Skip to content
Book Open access

MotRNA: Encoding RNA Motifs via Explicit N-gram Memory

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 14 references

TL;DR

Crucially, the analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.

Abstract

RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling global dependencies within the raw sequence. To address this limitation, we propose MotRNA, a motif-aware RNA language model integrating an Explicit N-gram Memory mechanism. Unlike standard implicit neural modeling, MotRNA employs a conditional memory module to explicitly retrieve and fuse motif embeddings based on input k-mers. This mechanism provides direct access to strictly defined local patterns, complementing the global context modeled by self-attention. We validate MotRNA on large-scale RNA datasets. The model achieves a 9.6% absolute improvement in Masked Language Modeling (MLM) accuracy at a 30% masking ratio compared to state-of-the-art baselines. Notably, MotRNA maintains high robustness under extreme masking ratios, indicating effective biological signal reconstruction from sparse context. Crucially, our analysis reveals that the backbone encoder becomes self-contained post-training: the memory module facilitates the internalization of motif semantics into the model parameters, allowing the backbone to retain performance advantages even when the memory is detached during inference.

Read PDF

Similar papers

Open access Aug 2026

Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT

Motivation RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes. Results We present SPIRAL, a layer-wise SAE analysis of BiRNA-BERT. Independent SAEs at layers 0, 5, and 11 expand each 768-dimensional hidden state into 6,144 features while preserving model behaviour (explained variance above 0.99997; masked-language-model sequence recovery near 99.7%). Tokenizer-aware offset propagation aligns features to nucleotides: at layer 5, 44.3% of tested features are significantly associated with bpRNA secondary-structure classes (mean enrichment 1.61 ×), and all 1,237 eligible features with RNAcentral RNA types. Sparse profiles raise k-nearest-neighbour balanced accuracy from 0.328 to 0.359 over dense embeddings at layer 5. Availability and Implementation Source code is available at https://github.com/SadatHossain01/SPIRAL; the code, evaluation data, and trained SAE checkpoints are archived at https://doi.org/10.5281/zenodo.21891845. Contact mrahman@cse.buet.ac.bd Supplementary information Supplementary data are presented alongside the manuscript.

M. Hossain, MD. Roqunuzzaman Sojib, Md Toki Tahmid et al. · 0 citations
Open access Aug 2026

PARNET: A CLIP-SEQ-BASED FOUNDATION MODEL FOR RNA SEQUENCE REPRESENTATION LEARNING

RNA-binding proteins (RBPs) orchestrate a complex combinatorial regulatory “code” that governs RNA splicing, stability, localization, and translation. Learning the relationship between RNA sequences and these processes is a central challenge in genomics. Foundation models, notably RNA language models, have emerged as the dominant approach, learning general-purpose representations from unlabeled sequence at scale. While RNA language models have demonstrated impressive performance across a broad range of downstream tasks, they generally learn from sequence reconstruction objectives alone, lacking direct connections to the regulatory principles that govern RNA function. Here we introduce Parnet, an RNA foundation model trained directly and exclusively on experimental CLIP-seq data. Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence. This CLIP-seq pretraining strategy departs fundamentally from the masked-language-modeling paradigm, anchoring learned RNA representations directly in measured protein–RNA interactions rather than sequence statistics. Parnet substantially outperforms its single-task predecessor RBPNet in binding profile and motif recovery, generalizes to unseen cell types and iCLIP data, and recapitulates position-dependent splicing regulation. Frozen Parnet embeddings, without task-specific fine-tuning, match or exceed the performance of both task-specific tools, as well as larger self-supervised RNA and genomic language models across diverse downstream tasks, including RNA biotype classification, lncRNA chromatin localization, translational efficiency, splice-site recognition, intron retention, and non-coding variant effect prediction. Importantly, Parnet remains mechanistically interpretable, tracing predictions back to the specific RBPs and motifs that drive them. These results establish the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.

Lambert Moyon, Andreina Tirabassi, Artem Baranowskii et al. · 0 citations
#machine learning Preprint Aug 2026

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. Native 10K pretraining preserves strong reconstruction at 10,240 tokens and, in a controlled long-context benchmark, maintains strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct short-context extrapolation, but induces substantially greater distal representation diffusion. Frozen RNA-type evaluations show that RIBOSPAN learns state-of-the-art RNA representations, with a particularly clear advantage on long RNAs. Across downstream biological benchmarks, RIBOSPAN emerges as the strongest encoder-only RNA foundation model, achieving state-of-the-art performance in both full-transcript biological property prediction and zero-shot mutation-fitness modeling. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning, biological prediction, and full-transcript mRNA design.

Ziyuan Wang, Bohao Tang, Fei Zhang et al. · 0 citations
Open access Jul 2026

Shifu: an integrated framework for deep learning of RNA secondary structure

Deep learning has advanced RNA secondary-structure prediction by bypassing explicit energy rules to capture long-range dependencies, yet progress is limited less by model scale than by how structures are measured: single scores hide where and why models fail, and benchmark scores can reflect memorization of one dataset rather than genuine generalization. We address this with Shifu, a framework of three coupled parts. Shifu-Corpus is a leakage-audited dataset of 254123 sequences from six databases, with family-aware splits certified free of exact and near-duplicate leaks. The Shifu Trifecta scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number. Shifu-LMR, a family of compact RNA language models, serves as controlled experiments: changing the training corpus shifts accuracy by 0.13, and a 65-million-parameter model, Shifu-LMR-Nano, leads on correctness while running on a laptop. We release the dataset, code, and model backbones.

Gabriel Galvez, Quentin Vicens · 0 citations
Open access Aug 2026

RNA-Lexis: a probabilistic algorithm using a non-parametric segmentation logic to detect meaningful sequences in RNA

Abstract Deciphering sequence–function relationships in long non-coding RNAs (lncRNAs) remains challenging due to rapid evolutionary turnover and limited primary sequence conservation. Alignment-based approaches often fail to detect functional domains, and fixed-length k-mer models inadequately capture variable-length regulatory elements. Here, we introduce RNA-Lexis, a non-parametric statistical framework for unbiased discovery of candidate RNA sequence elements. RNA-Lexis applies segmentation based on local conditional probabilities to identify non-random sequence extensions, enabling detection of recurrent, variable-length motifs without prior biological assumptions. Conceptually analogous to language segmentation, the framework partitions continuous RNA sequences into statistically defined units (“xmotifs” and “cores”), providing an interpretable representation of sequence architecture. RNA-Lexis reconstructs the modular organization of well-characterized lncRNAs, including XIST and NORAD. In additional case studies, RNA-Lexis prioritized recurrent GC-rich elements in SNHG14 that were tested experimentally and shown to bind histones in RNA pulldown assays. RNA-Lexis also identified recurrent LINC01001 core motifs that overlap chromatin interaction patterns detected by GRID-seq. These analyses support the use of RNA-Lexis to nominate candidate sequence elements for functional follow-up, while biological function remains dependent on orthogonal experimental validation. RNA-Lexis provides a statistically grounded and interpretable framework for motif-level analysis of lncRNAs. Rather than directly inferring function, the method identifies recurrent sequence architecture and prioritizes candidate elements for mechanistic testing.

Haim Y. Bar, Amit Felach, A. Bester · 0 citations