The performance and flexibility of State set the stage for scaling the development of AI models of cell state, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments.
Abstract
While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.
ScPILOT learns a generative latent representation through discriminator‐assisted training and separates perturbation inference into cell‐level response estimation from observed contexts and query‐specific response transfer using latent optimal transport, a query‐conditioned framework for transferring responses to previ...
Jia-Liang Wang, Zi-Qi Liu, Zheng-Qiang Zhang et al.· Advancement of science· 0 citations
TranScouter is introduced, a lightweight encoderdecoder framework that represents perturbed genes using LLM-derived embeddings of their text summaries and represents biological conditions using transcriptomic profiles of control cells from the target condition.
Predicting the transcriptomic consequences of cellular perturbations is challenging due to the sparsity of single-cell transcriptional responses, heterogeneous perturbation efficiency, and the difficulty of identifying the small subset of genes that are truly differentially expressed after perturbation. Current approac...
Michele Calabrò, Patrick Sheehan, F. Cambuli et al.· bioRxiv· 0 citations
The results show that compact biological representations can support accurate and computationally efficient perturbation prediction, and highlight the importance of perturbation representations and population-construction procedures in low-data benchmarks.
Dewei Hu, Marc Pielies Avellí, L. J. Jensen et al.· bioRxiv· 0 citations
PerturbLDM, a latent-diffusion framework for conditional generation of single-cell transcriptional responses, is introduced, showing support for conditional response generation across data scales and biological settings.
Li-Shan Yu, Kang-Lin Hsieh, Y. Chu et al.· bioRxiv· 0 citations
Overall, scLDM provides a robust and biologically consistent strategy for in silico perturbation screening, and exhibits strong interpretability, as the learned perturbation embeddings show high functional alignment with known biological mechanisms.