Skip to content
Open access

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Sep 2026 · Cell · Vol 189 19, pp. 5932-5944.e6 · 1 citation · 57 references
Medicine

Abstract

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Read PDF

Similar papers

Open access Sep 2026

Gravlax: an annotation-independent molecular evidence archive for single-cell RNA-seq

A cell-by-gene count matrix is the artifact of a single-cell RNA-seq experiment that is most often stored, shared, and reanalyzed. It is the output of a computation whose inputs are the sequenced molecules and a gene annotation, and while the molecules never change, the annotation is revised continually. Once the matri...

Rob Patro · 0 citations
Open access Aug 2026

Ultrafast and reference-free sequence discovery in single-cell data.

Malva is presented, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts.

D. León-Periñán, Nikos Karaiskos, N. Rajewsky · 1 citation
Sep 2026

An alignment-last approach enables rapid transcriptomic biomarker discovery in large cohorts

Unsupervised read-level differential analysis recovers established lncRNA biomarkers; uncovers new prognostic transposable-element reads in adrenocortical carcinoma and sarcomas; and extracts signals even from reads that fail to align.

Erkan Narmanli, Alexandre Lanau, Mara Neacsu et al. · 0 citations
Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations
Book Open access Aug 2026

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.

Jiaxian Yan, Xi Fang, Jintao Zhu et al. · 0 citations
Open access Sep 2026

scOLAR: Ontology-Anchored Open-Set Annotation of Single-Cell RNA-seq Data

Single-cell RNA sequencing profiles cellular heterogeneity at atlas scale, making automated annotation essential. However, target datasets often contain novel cell types missing from incomplete references. We present scOLAR, an ontology-guided open-set framework that learns prototypes over the Cell Ontology and uses bo...

Yu-Qiao Liu, Siyu Yi, Hengchuang Yin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.