A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Abstract
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
Manoj M Wagle, Yongheng Wang, Soham Samanta et al.· bioRxiv· 0 citations
iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks, bridging the gap between computational inference and biological intuition.
Hongming Guo, Tingfang Wu, Wen-Zheng Wang et al.· PLoS Computational Biology· 0 citations
In the tumor microenvironment, cell's state is influenced by cell-cell interactions (CCIs) with neighboring cells in its niches. Identifying dysregulated CCIs that are associated with pathogenic process pinpoints targets for drug discovery. Imaging-based spatial transcriptomics and single-cell RNA sequencing provide, respectively, single-cell spatial information and transcriptome-wide measurements needed to study CCIs, but neither modality provides both. Existing spatial transcriptomics foundation models also cannot effectively learn from spatially resolved single-cell data with full-transcriptome coverage, explicitly infer the CCI mechanisms driving cell state-niche associations, or interpretable enough to support direct biological interpretations. Here, we present GITIII-scale, a hierarchical, interpretable pan-cancer spatial transcriptomics foundation model for TME representation learning that investigates cell state-niche associations and their underlying ligand-receptor (LR) signaling pathways. GITIII-scale uses transformers to model interactions between pairs of cells at defined spatial distances, an interpretable single-layer graph transformer without a feed-forward network to decompose how each gene in a receiver cell is influenced by each neighboring sender cell, and a graph transformer to generate cellular-neighborhood embeddings. Trained on our assembled pan-cancer database of specimen-matched scRNA-seq and imaging-based spatial transcriptomics datasets, GITIII-scale generated TME embeddings that recovered niche-associated state changes more accurately than existing spatial transcriptomics foundation models in cancer types unseen during training. A case study of an unseen breast cancer dataset further demonstrated the model's interpretability by identifying potentially drug-targetable LR pathways associated with endothelial overgrowth and tumorigenesis.
Xiaohui Xiao, Jia-Shu He, Shiyang Zhang et al.· 0 citations
A structure-aware interleaved-attention graph learning framework, termed IAGRN, is proposed for GRN inference from scRNA-seq data that interleaves topology-constrained local attention with distance-aware global attention, enabling effective integration of structural priors and long-range regulatory signals.
Yue Wang, Si-Cheng Tian, Dan Li· International Journal of Mol...· 0 citations
Motivation Single-cell RNA sequencing (scRNA-seq) has become an attractive tool for studying complex diseases, in which transient cell states affecting diverse cell populations characterise disease development and progression. However, due to data sparsity and disease heterogeneity analysis is often challenging. With recent advances in machine learning, two widely used approaches have emerged for learning cellular representations: large-scale foundation models and biological knowledge-guided methods. Despite their complementary strengths, there is currently no unified workflow for systematically comparing and integrating these approaches. Results Here, we present scRepresenter, an open-source workflow for computing, integrating, and validating cellular embeddings derived from foundation models and biological knowledge-guided methods in the context of complex diseases. It consists of two components: a command-line workflow that computes cellular embeddings and performs downstream analyses, and an interactive Shiny application for visualizing and comparing the computed embeddings. scRepresenter supports four categories of cellular representations: (1) expression-based, (2) knowledge-guided, (3) foundation model-derived, and (4) hybrid embeddings that combine foundation model-derived representations with knowledge-guided representations. This approach takes a cell-by-gene count matrix as input and outputs an integrated object containing the computed embeddings. Then, this object can be uploaded into our interactive Shiny application to compare different embeddings. Availability The workflow is available at https://github.com/GuilhermePocas/scRepresenter Contact AL291@cam.ac.uk; MA2129@cam.ac.uk
Guilherme Pocas, Muhammad Umar, Oliver Davis et al.· bioRxiv· 0 citations
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
Yuhao Wang, Zelin Zang, Yuxuan Liu et al.· 0 citations