Skip to content
Preprint

Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models

Aug 2026 · 0 citations · 14 references
Biology

TL;DR

A representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks shows that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

Abstract

Genomic foundation models are increasingly reused as frozen feature extractors for downstream sequence prediction, offering a compute-efficient alternative to full fine-tuning. However, it remains unclear when biological information encoded by these models is accessible without task-specific adaptation. We present a representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks. We evaluate DNABERT-2, Nucleotide Transformer, HyenaDNA, GENERATOR-v2, and Omni-DNA under unified frozen-probing protocols, while separating diagnostic readout analyses from validation-selected checks. Our results reveal a consistent task-dependent pattern: frozen probes recover 95-100 % of fine-tuned performance on promoter tasks, but average splice-site recovery drops to 60-88 %. Frozen embeddings are also competitive on broad Genomic Benchmark tasks such as coding-region and species-discrimination classification, but show larger gaps on some regulatory and OCR tasks. Layer-wise probing, in-silico mutagenesis, variant-effect prediction, and embedding geometry show that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

View source

Similar papers

Preprint Jul 2026

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

A framework that combines sparse dictionary learning with causal intervention to extract, validate, and causally test interpretable features in genomic foundation models is introduced, and a reusable standard for interpretability claims in genomic deep learning is provided.

Sarwan Ali · 0 citations
Preprint Jul 2026

Pre-Registered External Evaluation Yields a Consistent Partial-Replication Category across Three Transcriptomic Foundation Models

A pre-registered, final-test-once evaluation framework that locks the outcome rule, seeds, and target-gene-grouped splits before any test data are seen, and scores each frozen representation against a strong expression baseline, a matched-capacity Gaussian control, and a within-split row-identity (shuffle) control.

Mehrdad Shoeibi, Niloofar Yousefi · 0 citations
Preprint Aug 2026

GenomeHarness: Harnessing Al Agents for Reliable Adaptation of Genome Language Models

GenomeHarness, an agentic harness for adapting genome language models through controlled search over fine-tuning recipes, improves mean test MCC in 47 settings, and shows gains on Genomic Benchmarks and on tasks where the root recipe is unstable or poorly matched.

Weicai Long, Yusen Hou, Houcheng Su et al. · 0 citations
Review Jul 2026

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.

Mahinaz A. Mashhour, Manal Abdel Wahed, Mai S. Mabrouk · 0 citations
Open access Jul 2026

xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction.

Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA-RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.

Shumin Li, Ruibang Luo, Yuanhua Huang · 0 citations
Open access Aug 2026

Pretraining Enhances Megabase-Scale Gene Expression Prediction with GeneUnet

GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2 is introduced, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data.

Ning Sun, William de Vazelhes, Pan Li et al. · 0 citations