Aug 2026· Blood Science· Vol 8, pp. e00306 - e00306· 0 citations· 28 references
Medicine
TL;DR
Perturb-Audit, a diagnostic and denoising framework that integrates molecule-level collision auditing with statistical background suppression, supports an audit-first strategy to improve assignment fidelity and biological interpretability in single-cell CRISPR screens.
Abstract
Perturb-seq enables high-throughput linkage of CRISPR perturbations to single-cell transcriptomic phenotypes; however, inference quality depends on accurate single-guide RNA (sgRNA) assignment. In 10x Genomics–based single-cell workflows, assignment can be distorted by ambient RNA, overloaded droplets, and amplification artifacts, including cross-library polymerase chain reaction (PCR) chimeras. We demonstrate that standard Cell Ranger processing—with independent correction of gene expression and CRISPR libraries—does not explicitly resolve cross-library molecular collisions, in which a single-cell barcode–unique molecular identifier (CBC–UMI) pair is assigned to discordant features. To address this limitation, we developed Perturb-Audit, a diagnostic and denoising framework that integrates molecule-level collision auditing with statistical background suppression. Across Perturb-seq datasets of T-cell exhaustion, targeted collision removal provides high-specificity cleanup, whereas global denoising with CellBender yields broader improvements in assignment quality and phenotypic separation. Improved assignment fidelity increases detectable perturbation effect sizes and enables the recovery of biologically relevant immune cell signals. Applying this approach, we recapitulated the known Klf2-deficient phenotype in antiviral CD8+ T-cell Perturb-seq data. Furthermore, we found that suppression of Eomes triggers an exhaustion-biased shift, whereas Tox deficiency promotes effector-like differentiation. Collectively, these findings support an audit-first strategy to improve assignment fidelity and biological interpretability in single-cell CRISPR screens.
Perturb-seq enables pooled genetic screens with rich single-cell profiling readouts, but genome-scale profiling remains costly and may not be associated with other established functional characteristics. Moreover, as screens grow in size and complexity, interpreting the resulting data comprehensively is challenging and slow. Here, we introduce Perturb-seq with Marker Enrichment (Perturb-ME), which combines genome-scale CRISPR screening, phenotype-based enrichment and multimodal single-cell profiling. Applied to MHC-I cell surface protein expression in melanoma, Perturb-ME profiled HLA-low and HLA-high cells with matched RNA, surface-protein and guide measurements. A regulatory model with 221 impactful regulators affecting 1,998 responsive genes recovered seven coherent co-functional regulatory modules governing nine gene programs, including the canonical IFNγ-MHC-I axis regulating an antigen-presentation and interferon-response program. Agentic interpretation of the entire model with an AI co-scientist linked additional modules to trafficking, proteostasis and chromatin regulation. Perturb-ME, along with agentic interpretation, provide a scalable framework for comprehensive functional discovery from phenotype-enriched genetic screens.
Hanchen Wang, Jiacheng Gu, Chris J. Frangieh et al.· bioRxiv· 0 citations
Accurate detection of off-target activity in primary human cells is crucial for ensuring the safety of gene therapies, yet existing methods often lack sufficient sensitivity. To address this limitation, we develop Tracking-seq2, an advanced technology that integrates exogenous 5′ → 3′ exonuclease treatment and non-homologous end joining (NHEJ) pathway inhibitors with the original Tracking-seq. Tracking-seq2 exhibits enhanced sensitivity in profiling off-target sites of diverse genome editors—including Cas9, Cas12a, cytosine base editors (CBEs), adenine base editors (ABEs), and prime editors (PEs). Critically, Tracking-seq2 is directly applicable to clinically relevant primary human cell types, such as T cells and CD34
+
hematopoietic stem and progenitor cells (HSPCs). Furthermore, our findings reveal that genomic variations drive distinct off-target heterogeneity across different individuals, highlighting the necessity for personalized safety assessment in clinical genome editing applications. Tracking-seq2 provides a robust platform for sensitive off-target detection in primary cells, with sensitivity comparable to or exceeding current state-of-the-art methods.
Single-cell RNA sequencing (scRNA-seq) pipelines rely on the assumption that sequencing reads possess correct structural architecture, a premise we show is incomplete. Standard quantification tools treat errors exclusively as base mismatches, failing to identify structural aberrations arising from off-target priming or nonspecific amplification. We demonstrate that these pervasive artifacts, reads lacking essential anchor motifs like poly(T) tracts, linkers, or template-switching oligos, generate large numbers of spurious barcodes, artificially inflate cell counts, and substantially affect biological interpretation. To resolve this, we developed Pattern-Filter, a universal preprocessing tool that systematically validates read integrity before alignment. It functions by detecting platform-specific anchor sequences and applying strict base-composition filtering to ensure barcodes and UMIs contain only canonical nucleotides. When applied across diverse platforms, including 10x Genomics, Drop-seq, BD Rhapsody, and SPLiT-seq, Pattern-Filter systematically removes 2%-18% of total reads yet reduces spurious barcode diversity by up to 80%. This asymmetric reduction confirms that a small fraction of invalid reads drives the majority of technical noise, compromising cluster stability. Consequently, this targeted removal enhances data reproducibility and recovers biologically relevant cell types, such as dopaminergic neurons in mouse striatum, which were previously obscured by artifact-induced noise. These findings establish structural validation as an essential prerequisite for analysis, positioning Pattern-Filter as a useful standard for ensuring molecular fidelity and reliable biological discovery in single-cell transcriptomics.
Qiang Su, Xiaoming Zhou, Yi Long et al.· Genome Research· 0 citations
Perturb-seq measures transcriptomic responses to genetic perturbations at scale, but conventional designs that enrich for one guide RNA per cell remain resource-intensive. Standard analyses discard cells carrying multiple guides, further limiting the usable yield from each experiment. Here, we characterize how incorporating these guide multiplets affects signal recovery, information loss, and cost reduction. At the highest guide burden, cells showed increased stress and suppressed cell-cycle progression. We develop PerturbMatch, a scalable statistical framework to analyze guide multiplets. Among different classes of guide multiplets, doublets and triplets recovered perturbation responses more accurately than higher-order multiplets. Across three 5000-gene Perturb-seq screens with increasing guide loading, per-cell costs decreased by up to 81% while information loss remained within 1.5-fold of the loss observed between technical replicates. In existing genome-wide Perturb-seq data, incorporating previously discarded guide multiplets increased usable cell numbers and improved statistical power. Compared with a singlet holdout set, adding guide multiplets moved signal recovery closer to the theoretical expected reproducibility. Overall, we recommend a design that intentionally includes single-guide cells, guide doublets, and guide triplets to improve cost efficiency while preserving signal recovery.
Jake Yeung, Jenille Tan, Liang Wang et al.· bioRxiv· 0 citations
Deleting a gene token from a cell’s input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformer’s perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation. Motivation Foundation-model in silico perturbation could predict perturbation effects when matched experimental data are unavailable. However, in zero-shot settings, embedding-derived responses may reflect gene identity, universal responsiveness, tokenization limits, library-size artifacts, or circular state scoring rather than biological knockout effects. We therefore developed a reusable confound-diagnostic framework that applies matched controls to test whether native perturbation readouts contain information beyond these confounds and warrant biological interpretation.
An in-cell method that sensitively detects rare off-target edits, benchmarks performance across editors, and improves risk assessment in therapeutic cells is presented, establishing UNCOVERseq as a robust framework for informed off-target risk assessment in translational gene-editing systems.
Kyle J. Kinney, Kun Jia, He Zhang et al.· Nature Communications· 3 citations