Skip to content
Open access

Benchmarking cell type annotation in spatial transcriptomics: resolving cellular hierarchies, biological fidelity, and dynamic cell states

Aug 2026 · Research Square · 0 citations · 45 references
Medicine

TL;DR

A comprehensive, context-specific guide to current annotation strategies for spatial transcriptomics is presented and open-set recognition of reference-absent cell states, adaptive incorporation of spatial context, and improved resolution of rare and transitional cell identities are identified as central priorities for the next generation of annotation methods.

Abstract

Spatial transcriptomics enables the quantification of gene expression within its native tissue context, providing unprecedented insight into tissue architecture, cellular ecosystems, and local cell–cell interactions at regional and single-cell resolution. Accurate cell type annotation is a critical prerequisite for interpreting these data and is often the first and most essential step in downstream analysis. Despite rapid advances in computational methods, cell type annotation remains challenging and frequently requires extensive expert-driven manual curation based on marker-gene expression, spatial context, and prior biological knowledge. While early approaches relied primarily on transcriptional similarity, newer methods increasingly incorporate spatial information, histological features, and multimodal data to improve annotation accuracy. Nevertheless, reliable annotation remains difficult when biological interpretation requires fine-grained subtype resolution, particularly for platforms with limited gene panels, tissues undergoing dynamic cellular state transitions, and studies in which reference and query datasets differ substantially in biological context or technical modality. Here, we present a systematic benchmark of 20 state-of-the-art annotation methods across four spatial transcriptomics technologies and six biologically and technically distinct benchmarking scenarios spanning diverse technologies, experimental conditions, cell numbers, and gene panel sizes. Importantly, all benchmark datasets contain expert-curated cell type labels, including well-resolved cell populations and subtype annotations, providing high-quality biological ground truth for evaluation. The benchmark encompasses both reference-based and reference-free methods representing a broad range of computational frameworks. Performance was assessed using conventional classification metrics, including accuracy and F1-based measures, together with structure-aware metrics that evaluate both cell-level annotation accuracy and preservation of higher-order biological organization. Across datasets, annotation performance varied substantially according to tissue context, reference–query similarity, and annotation granularity. Fine-grained subtype annotation and recovery of rare cell populations remained challenging for many methods, particularly in datasets capturing injury, repair, developmental, and regenerative processes characterized by continuous cellular state transitions. Notably, high classification accuracy did not necessarily correspond to preservation of global cellular relationships or biologically coherent downstream pathway and gene-set enrichment analyses. Overall, scANVI, Seurat, and TACCO consistently ranked among the top-performing methods, although the best choice varied by context: methods effective for within-platform reference transfer or canonical, well-separated cell types were not necessarily the strongest under cross-platform, cross-developmental-stage, or disease-dynamic transfer. Together, our results provide a comprehensive, context-specific guide to current annotation strategies for spatial transcriptomics and identify open-set recognition of reference-absent cell states, adaptive incorporation of spatial context, and improved resolution of rare and transitional cell identities as central priorities for the next generation of annotation methods.

Read PDF

Similar papers

Aug 2026

Accurate reconstruction of spatial cell type maps and characterization of domain-specific functions based on a gene-aware heterogeneous network.

Spatial transcriptomics (ST) profiles gene expression with spatial context, but most platforms capture multicellular spots containing mixed cell types, making accurate deconvolution essential. Existing reference-based methods using scRNA-seq often ignore spatial dependency and gene-level contribution, yielding fragment...

Zi-Lin Li, Zhao-Yang Huang, Yan Li et al. · 0 citations
Review Sep 2026

Single-cell and spatial transcriptomics inform mechanistic physiology in non-model animals

This review provides a physiology-centered blueprint for applying single-cell RNA sequencing, single-nucleus RNA sequencing, and spatial transcriptomics to non-model species and critically evaluates dissociation and preservation bias, genome annotation, seasonal and ecological variation, biological replication, pseudor...

Adnan Amin, W. Zaman · 0 citations
Aug 2026

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

This protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation by emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints.

Hua-Lin Wang, Wei-Jia Chen, Yan Wu et al. · 0 citations
Open access Aug 2026

Systematic evaluation of spatial transcriptomic annotation methods reveals conserved tumor microenvironment programs in NSCLC

Introduction Spatial transcriptomics (ST) enables high-resolution mapping of cellular heterogeneity in tumor microenvironments, but downstream cell-type annotation remains highly method-dependent. Methods Here, we systematically evaluated four representative annotation frameworks—Seurat, CARD, RCTD, and SPOTlight—acros...

Mingke Wu, Yi Ting Zhou, Ding-Jie Xu · 0 citations
Open access Sep 2026

Global tree encoding of atlas-scale single-cell genomics

MILK is presented, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations that establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell gen...

Brett Kiyota, Chaehyeon Lee, Hao-Yang Yao et al. · 0 citations
Open access Aug 2026

GRIDGENE: Guided Region Identification based on Density of GENEs – a transcript density-based approach to characterize tissues by spatial transcriptomics

Spatial omics brought unprecedented power to study biological processes within tissues while preserving spatial context and morphology. Most spatial proteomics and transcriptomics analysis methods are cell-centric, relying on cell segmentation to identify and characterize individual cells before downstream tasks. How...

A. M. Sequeira, M. Ijsselsteijn, M. Rocha et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.