Skip to content
#small language model Open access

CytoGate-Bench: an LLM benchmark for cross-panel cell gating in cytometry

Aug 2026 · bioRxiv · 0 citations · 5 references
Biology

TL;DR

This work introduces CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models, and contributes a public benchmark that tests precisely that ability across 11 human cohorts.

Abstract

In cytometry, the workhorse single-cell technology of clinical immunology, every study defines its own antibody panel and cell-type vocabulary, so a classifier trained on one cannot annotate the next. Immunologists instead annotate by manual gating, splitting one parent population at a time on a two-marker plot, down an expert-defined hierarchy. We introduce CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models. It comprises 23,646 expert-annotated instances re-curated from 11 public flow- and mass-cytometry cohorts spanning eight marker panels. Across six open- and closed-weight backbones, the strongest formulation draws one rectangular gate per candidate and falls within the range of trained, panel-specialized baselines. It degrades less under distribution shift. Walking the hierarchy stepwise outperforms predicting every cell type at once. Ablations trace the signal to the data distribution shape and curated marker priors. However, adding vision or a self-verification loop systematically tightens gates. THE BIGGER PICTURE Immunology laboratories worldwide profile blood and tissue with cytometry, an instrument family that measures dozens of protein markers on millions of individual cells. Before any biology can be read out, every cell must be assigned an identity, a step still dominated by manual “gating,” in which an expert draws boundaries on a sequence of two-marker plots, following a documented, hierarchical protocol. Automating this step has remained difficult because every study measures a different marker panel and names a different set of cell types, so conventional machine-learning models must be retrained for each new study. Large language models (LLMs) promise a different route, a single general-purpose model that reads the expert’s protocol and the data and makes each gating decision directly, with no study-specific training. This work contributes a public benchmark that tests precisely that ability across 11 human cohorts. The result is a statement of feasibility rather than superiority. Off-the-shelf models already score in the range of study-specific trained models and tolerate the day-to-day variation that degrades them. That capability matters most for new or small studies, for which no labeled training data exists. Walking the expert’s hierarchy one decision at a time also outperforms asking the model to name every cell type in a single pass, evidence that the structure of expert practice matters more than the scale of the question. These results come from a deliberately minimal setup, untuned models drawing simple rectangular gates, so we read them as a floor rather than a ceiling. Cytometry-aware training, richer gate geometries, and better-calibrated visual feedback are open avenues, and the benchmark gives that progress a fixed yardstick. Sustained progress would give laboratories analysts that keep pace with evolving marker panels without retraining, while leaving a decision trail an immunologist can audit.

Read PDF

Similar papers

Preprint Aug 2026

CytoBERT: A Foundation Model for Cytometry Data

Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. However, cytometry data is highly heterogeneous and unstandardized due to experimental protocols and the choice of measured features. While machine learning methods hold the potential to gain deeper insights into cell biology, these challenges make them difficult to apply and transfer across studies. Recent advances in foundation models can alleviate these issues, but corresponding approaches are still scarce in this field. To address this, we provide CytoBERT, a publicly available, open-source, open-weight foundation model for single-cell cytometry data with variable marker panels. CytoBERT is pretrained in a self-supervised manner on a large-scale cytometry corpus (15 human datasets with heterogeneous marker panels and more than 50 million cells) curated through marker standardization, enabling it to learn transferable inter-marker relationships within cells. Fine-tuning CytoBERT for sample-level classification demonstrates that transfer learning across heterogeneous cytometry datasets is feasible, providing a starting point for scalable, generalizable cytometry analysis. Code is available at GitHub.

Syed Abdul Haseeb Qadri, Bjarne C. Hiller, F. Blanke et al. · 0 citations
Open access Jul 2026

A Label-Free Multi-Metric Pipeline for Benchmarking Single-Cell RNA-Sequencing Clustering and Testing the Reproducibility of Cell-Type Heterogeneity

A discovered sub-population from single-cell transcriptomic data is only meaningful if it is reproducible, yet clustering is usually done with one method on one embedding and rarely tested. We present a label- free, multi-metric pipeline that reframes clustering as an auditable, methods-blind decision and separates two notions of stability that are commonly conflated: reproducibility under cell resampling (bootstrap) and reproducibility under re-embedding (retraining the representation). The pipeline evaluates seven clustering configurations across cluster counts using five non-redundant quality metrics. As a whole- dataset control on a mouse retinal atlas, it recovers an eight-cell-type annotation at 96.3% accuracy (adjusted Rand index, ARI = 0.91) without labels. We then validate the discovery mode on two cell types with opposite ground truth. On bipolar cells, which have well-established subtypes, the pipeline accepts the sub-structure: across-embedding reproducibility rises with cluster number to a high plateau (mean pairwise ARI ∼0.93 near the ∼15 known bipolar subtypes), with quality metrics improving in parallel. On rod photoreceptors, treated as homogeneous, it rejects over-clustering: the metric-selected partition passes a bootstrap-stability check but is not reproducible when the embedding is retrained (mean pairwise ARI = 0.69), and the metrics do not improve with cluster number. On synthetic data, the test recovers real structure down to a 5% subpopulation while rejecting null data (high sensitivity and specificity). Bootstrap stability alone is therefore insufficient evidence for sub-population; the across-embedding test discriminates real sub-structure from over-clustering and applies to any cell type as a reproducible alternative to single-method, single-embedding clustering.

Zachary Yousef, Jonah Simone, Dylan Klein et al. · 0 citations
Open access Jul 2026

Resolving Immune Lineage and Cell-State Heterogeneity in Human PBMCs via Mass Spectrometry-Based Single-Cell Proteomics

Single-cell proteomics (SCP) currently lacks validated benchmarking standards, and cell annotation often relies on transcriptomic proxies. Unsupervised clustering offers a proxy-free alternative, but its success depends on biological signal outweighing technical variation. In homogeneous samples this is achievable, but in heterogeneous populations, where closely related cell types differ only subtly, technical variation can dominate the clustering and obscure the biology needed for annotation. To address this, we developed an integrated experimental and computational pipeline for protein-level cell annotation and applied it to human PBMCs as an immune-cell test case. We isolated T cells, B cells, monocytes, and NK cells by negative-selection sorting to build a high-fidelity reference. In parallel, unsorted PBMCs from the same donor were processed on a cellenONE and acquired using label-free DIA on an Orbitrap Astral Zoom. Using the labeled reference dataset, we systematically benchmarked normalization, imputation, and clustering methods to assess their effect on cell-type separation. Unsupervised analysis resolved functional subpopulations within each lineage, and a probabilistic SCP classifier trained on these annotations identified the corresponding cell types and states in the unsorted PBMC fraction, validating the pipeline on unenriched, heterogeneous samples. Together, this work delivers an analytically benchmarked SCP workflow that resolves immune lineage and cell-state heterogeneity in human PBMCs and provides a classifier-ready, protein-level reference for immune-cell assignment.

Samantha A. O’Connor, Romell B. Gletten, Ritin Sharma et al. · 0 citations
Preprint Aug 2026

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information and the transferability of such information across tissue sections, datasets, and anatomical organs. We introduce CellPath-Bench, a cellular-resolution benchmark that evaluates frozen PFMs themselves. Following quality control of 52 candidate Xenium datasets, we construct a panel of 25 spatially aligned H\&E--Xenium tissue sections spanning 11 organs and 7,079,283 cells, harmonized into fine- and coarse-grained taxonomies. CellPath-Bench samples frozen WSI feature maps at registered nuclear coordinates and evaluates them using standardized multiclass linear probes. Cell Representation Advantage (CRA) measures the within-section advantage of nucleus-anchored representations over patch-level mean pooling, while Cell Representation Transferability (CRT) characterizes the generalization of cell-type decodability across tissue sections, datasets, and organs. We benchmark 30 pathology-specific and general-purpose foundation models through 304,920 runs across spatial readouts, magnifications, taxonomic granularities, and evaluation protocols. The results reveal substantial model-dependent differences in cell-type decodability and its cross-domain generalization, yielding distinct multidimensional capability profiles. CellPath-Bench provides a standardized framework for auditing cellular information in frozen PFM representations.

Bokai Zhao, Yiyang Zhang, Hanqing Chao et al. · 0 citations
Jul 2026

A 61-Parameter CyTOF Panel for Comprehensive Profiling of Human PBMC to Characterize Activation, Differentiation, Checkpoints and Cytokines 2259198

High-parameter cytometric analysis enables discoveries of potential therapeutic targets and informs disease prognoses in translational and clinical research. CyTOF™ technology is a single-cell analysis platform that uses metal-tagged antibodies to resolve 50-plus markers in a single tube. Notably, antibody cocktails and stained samples can be frozen for later use and acquisition, or barcoded and pooled together, minimizing technical variation. Unlike fluorescence-based cytometry, spectral unmixing and single-stain controls are not required. Therefore, it is uniquely possible with CyTOF technology to rapidly design high-parameter panels and use intracellular markers to gain functional insights. The goal of this study was to design a 61-parameter CyTOF panel for in-depth functional immune profiling of human PBMC. The panel contains lineage markers for major immune cell subsets and a diverse array of phenotyping T cell targets focused on activation, differentiation, checkpoint and cytokines and can be used to identify over 60 cell populations. Untreated and stimulated PBMC were barcoded, pooled and stained. Samples were frozen and acquired on a later day using a CyTOF XT PRO system. High-dimensional analysis of stimulated PBMC revealed striking cross-lineage immuno-functional diversity at the single-cell level. The expression of over 20 markers spanning the functional landscape of activation, checkpoints and cytokines was revealed across effector, memory, cytotoxic, regulatory and exhausted immune cell populations. CyTOF systems enable the highest number of simultaneous measurements in a single panel, allowing for wide immune coverage and high resolution of intracellular targets to interrogate functional potential. Overall, studying functional immunology using CyTOF technology can elucidate the complex nature of immune responses to provide an understanding of how they relate to disease and treatments. For Research Use Only. Not for use in diagnostic procedures., n/a Immune Mechanisms of Human Disease (HUM)

Michael J. Cohen, Stephen Li, Lauren J. Tracey et al. · 0 citations
Open access Jul 2026

scINTILLA: Single-Cell Integrated Inference, Labelling, and Landscape Analysis for Cell-Type Annotation Quality Assessment

Single-cell RNA sequencing has enabled the construction of comprehensive cell atlases, yet the quality and coherence of the cell-type annotations within these atlases remain largely unexamined. When a label is applied to a transcriptionally heterogeneous population, the downstream analyses that depend on it, and automated label transfer in particular, become unreliable. We present scINTILLA (Single-Cell Integrated Inference, Labelling, and Landscape Analysis), a computational framework that combines supervised and unsupervised machine learning to score the learnability and internal consistency of cell-type labels in single-cell datasets. The unsupervised arm benchmarks a broad panel of clustering algorithms and derives a neighbourhood confusion score for every cell, whilst the supervised arm trains up to twelve classifiers and extracts prediction agreement, entropy, and confidence. These signals are normalised and aggregated into a single composite score per cell type, where a low score flags label ambiguity or concealed heterogeneity. As a by-product, scIN-TILLA also reports which clustering and classification algorithms perform best on a given dataset, offering practical guidance for downstream label transfer. We applied it to five Human Cell Atlas datasets spanning the adult brain, lung, eye, and two organoid atlases, and recovered clear differences in the learnability and internal consistency of annotations across atlases that were not driven by the number of annotated cell types. Focused re-analysis of lowscoring populations in the lung and endoderm-organoid atlases resolved biologically coherent sub-populations, in some cases with context-specific enrichment, much of it recovered from cells that had been assigned broad or catch-all labels. scINTILLA is advisory rather than prescriptive, guiding principled, data-driven re-annotation at atlas scale.

Sina Kanannejad, Noemi Bongiorni, Elisa Nordera et al. · 0 citations

Related blog posts