Skip to content
Open access

Nonparametric kernel-based detection of spatially variable genes with adaptive shrinkage and scalable multi-sample inference

Aug 2026 · bioRxiv · 0 citations · 34 references
Biology

TL;DR

CytoKspace is introduced, a nonparametric framework that combines a sparse exponential kernel constructed from nearest-neighbor graphs with a quadratic-form test statistic and adaptive permutation testing that demonstrates competitive sensitivity, well-calibrated false positive rates, and practical computational requirements compared to existing methods.

Abstract

Identifying spatially variable genes (SVGs), genes whose expression varies coherently across tissue space, is a central analytic goal in spatially resolved transcriptomics. Current methods rank spatially variable genes using either significance probabilities from parametric models or effect sizes such as the proportion of spatial variance, but parametric approaches impose distributional assumptions, such as Gaussian processes or negative binomial models, that may be violated for sparse or zero-inflated data. Furthermore, most detection tools cannot jointly model multiple biological replicates, and no existing framework provides both nonparametric significance probabilities and stabilized effect-size estimates with formal uncertainty quantification. Here, we introduce CytoKspace, a nonparametric framework that combines a sparse exponential kernel constructed from nearest-neighbor graphs with a quadratic-form test statistic and adaptive permutation testing. CytoKspace employs a multi-stage adaptive permutation schedule that yields substantial computational savings over fixed-permutation baselines, an adaptive shrinkage layer built on empirical Bayes estimation that stabilizes raw spatial effect sizes and provides posterior estimates with local false sign rates, and a scalable multi-sample extension via Fisher combination of significance probabilities and inverse-variance-weighted meta-analysis that accommodates studies with multiple biological replicates. In extensive simulations across a broad range of sample sizes, gene counts, spatially variable gene fractions, and effect sizes, as well as in applications to two real datasets from the Visium and seqFISH platforms, CytoKspace demonstrates competitive sensitivity, well-calibrated false positive rates, and practical computational requirements compared to existing methods. A software implementation of our method is freely available at https://github.com/Ghoshlab/CytoKspace.

Read PDF

Similar papers

Open access Sep 2026

Structured Spike-and-Slab Variational Bayes for High-Dimensional Non-Normal Generalized Linear Mixed Models

High-dimensional non-normal longitudinal data are ubiquitous across fields such as genomics, biomedicine, microbiome research, and the social sciences. Such data often combine non-normal responses, within-subject dependence and sparse population effects. We develop a structured variational Bayesian procedure: variation...

Jie-Yi Yi, Ying Wu, Yun-Qi Zhang · 0 citations
Open access Sep 2026

Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes

Gaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. Covariance non-stationarity has long been recognized in spatial statistics as an important feature of spatial data, yet it has received little attention in spatial transcriptomics. We show that this omission...

P. Velidi, Zheng-Hong Wei, F. Nathoo · 0 citations
Preprint Aug 2026

Spatially orthogonal factor models for spatial transcriptomics and remote sensing data

Principal component analyses are often applied to spatial data towards inference on latent modes of spatial variation. These analyses are widespread across domains including spatial transcriptomics and environmental sciences, where the modes of spatial variation are represented by corresponding factors of gene expressi...

Cun-Ha Dan, Lukas M. Weber, M. Friedl et al. · 0 citations
Preprint Sep 2026

Covariate-localized False Discovery Rates

Several procedures for estimating and thresholding the local false discovery rate are introduced, and it is shown that this holds for fixed and randomized hypothesis labels, indicating that the proposed methods perform well under both frequentist and Bayesian interpretations of multiple testing.

Jonathan Lin, S. Tokdar · 0 citations
Preprint Aug 2026

A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment

Simulations show that joint estimation is useful when regression surfaces change abruptly across spatial boundaries, including settings with nonlinear effects, unequal region sizes, preferential sampling, and spatially correlated errors.

Ke-Xuan Li, Wei-Dong Ma · 0 citations
Preprint Sep 2026

Mixture-based Nonparametric Estimation of Spatial Covariance Functions with Applications to HIV Key Population Size Estimation across Sub-Saharan Africa

Consistent data on the sizes of key populations, such as female sex workers (FSWs), are often scarce, particularly at the sub-national level. Accurate size estimates are critical to effectively allocate resources and achieve HIV targets. Since FSW population sizes may be spatially correlated across areas, models that a...

M. Siriwardana, Hyebin Song, Le Bao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.