Skip to content
Open access

LucaCell: a sequence-centric foundation model for cross-species single-cell analysis

Sep 2026 · bioRxiv · 0 citations · 41 references
Biology

TL;DR

LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations, is presented, showing that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictive tasks.

Abstract

Single-cell foundation models have transformed transcriptomic analysis, yet most rely on fixed gene identifiers that limit transfer across species and data types. Here we present LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations. Gene expression is discretized into bins and modeled with a Transformer encoder, enabling sequence-informed cell representation without a fixed gene-ID vocabulary. Pre-training on 85 million human and mouse single cells, LucaCell is evaluated on human, mouse and lemur gene expression profiles, human chromatin accessibility data, unaligned reads from more than 50 prokaryotic taxa, and five influenza A virus genomes. LucaCell enables manual-mapping-free cross-species cell type annotation and an alignment-free microbial embedding framework that simultaneously distinguishes bacterial species identity and intra-species physiological states. It also improves gene expression reconstruction by incorporating donor-specific exonic SNP information into mRNA sequence embeddings, and predicts cellular viral load across influenza A virus strains while highlighting infection-like transcriptional states in mock-infected cells. These results show that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictive tasks.

Read PDF

Similar papers

Open access Sep 2026

scMaize: A Single-Cell Foundation Model and Integrated Atlas for Maize

Single-cell transcriptomics has resolved cell-type-specific gene expression in plants, yet maize still lacks an integrated reference and species-specific foundation models. We present scMaize, combining scMaizeAtlas, an integrated atlas of 385,675 cells from 20 projects and 66 samples across seven tissues with hierarch...

Qian Cheng, Ying Zhang, Tianhao Wu et al. · 0 citations
Open access Aug 2026

Ultrafast and reference-free sequence discovery in single-cell data.

Malva is presented, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts.

D. León-Periñán, Nikos Karaiskos, N. Rajewsky · 1 citation
Open access Sep 2026

Setting the SCENE for Interpretable Cell–Gene Embeddings in Single-Cell RNA-seq

Single-cell RNA sequencing measures cellular states at high resolution, but sparse high-dimensional count data remain difficult to model interpretably. We introduce the Single-Cell Euclidean Network Embedding (SCENE), a probabilistic latent-distance model that jointly embeds cells and genes from Unique Molecular Identi...

O. Møberg, M. Petersen, Tue Herlau et al. · 0 citations
Preprint Aug 2026

Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views

The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies but does not directly optimize whole-cell representations, which are essential for many...

Jiaqi Xiong, Yun-Tao Hu, Yu Zheng et al. · 0 citations
Open access Aug 2026

AdaGeneBudget: Cell-Adaptive Gene-Token Allocation for Efficient Single-Cell Foundation Models

AdaGeneBudget is introduced, a training-free gene-token selection method that combines each gene’s expression with reference-derived inverse detection frequency and retains the shortest ranked prefix that captures a target fraction of the cell’s expression-specificity score mass, which establishes biologically informed...

Dohee Kim, Uiwon Hwang · 0 citations
Preprint Aug 2026

CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific techni...

Hai-Ping Liu, Qian Zhao, Lijing Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.