Skip to content
Open access

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Abstract

Public transcriptomic repositories contain millions of samples, yet their large-scale reuse is hindered by heterogeneous and inconsistently reported metadata. In the Gene Expression Omnibus (GEO), key biological information is often distributed across study- and sample-level records, requiring context-dependent interpretation. Here we present GEOMeta, a large language model (LLM)-based multi-stage workflow with task-specialized agents for automated GEO metadata curation. The pipeline separates metadata retrieval, task-specific information extraction, field standardization, ontology mapping and quality control. Using GEOMeta, we generated standardized annotations for approximately 600,000 human bulk RNA-seq samples. To demonstrate its utility, we benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings. We further prospectively annotated newly submitted GEO studies and evaluated 22 frontier LLMs. Recent open-source Flash models achieved annotation quality comparable to leading reasoning models while reducing costs by an order of magnitude. GEOMeta provides a scalable resource and reproducible framework for metadata curation.

Read PDF

Similar papers

Review Open access Aug 2026

Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.

P. Pavlidis, B. O. Mancarci, A. Mãximo et al. · 0 citations
Open access Aug 2026

Automated generation of a gene perturbation transcriptomic atlas using large language models

Public transcriptomic repositories contain thousands of gene perturbation experiments, a valuable resource for understanding gene function, but perturbation metadata are not structured, which blocks systematic reuse. Existing perturbation atlases depend on expert manual curation, so they are costly to maintain and infrequently updated, while automated grouping approaches neither identify which samples form the perturbation arm nor recover the perturbed gene. Here we develop an automated pipeline that uses large language models to find single-gene perturbation experiments in NCBI-GEO and reconstruct their case-control sample groupings, along with the perturbed gene, perturbation type and cell line as structured, ontology-normalised fields. We manually curated 3,300 GEO experiments with sample-level case-control assignments and release these as an open benchmark (2,400 training, 600 validation, 300 temporally held-out test). Reasoning models and task-specific finetuning substantially improved identification of valid perturbation groups, with the best model reaching precision 0.925 and recall 0.836 on the test set. Applied at scale, the pipeline generated an atlas of 6,802 gene perturbation expression signatures from 4,453 GEO experiments, covering 2,907 uniquely perturbed genes. An R package, perturbMatch, supports exploration of the atlas and querying of user-supplied expression signatures against it using similarity scoring, so users can identify experiments that recapitulate a transcriptional state of interest.

J. Soul, D. Young · 0 citations
Open access Jul 2026

LLM-powered Functional Gene Set Summarization with genesetGPT

GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.

Jack R. Leary, Samantha Pattey, Rhonda L. Bacher · 0 citations
Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations
Open access Aug 2026

Large language models enhance annotation of enzymes in metagenomes

FEDKEA, an enzyme annotation tool leveraging protein language models, and a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses are designed.

Lei Zheng, Bowen Li, Siqi Xu et al. · 0 citations
Open access Jul 2026

AtlasLens: Metadata-centric exploration and analysis of single-cell atlases

AtlasLens is developed, an open-source R/Shiny application for interactive exploration of scRNA-seq datasets and integrated cellular atlases that integrates interactive visualization, differential expression analysis, Gene Ontology enrichment with redundancy reduction, temporal expression analysis, and context-dependent gene function profiling through GeneCOCOA.

Shamim Ashrafiyan, Iaroslav Kosaretskii, Marcel H. Schulz · 0 citations