Aug 2026· PLoS Computational Biology· Vol 22 8, pp.
e1014727
· 0 citations· 69 references
Medicine
TL;DR
iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks, bridging the gap between computational inference and biological intuition.
Abstract
Precise resolution of cellular heterogeneity within complex tissues is fundamental to deciphering disease etiologies from bulk transcriptomic profiles. While computational deconvolution offers a scalable alternative, current deep learning methods predominantly operate as "black boxes," neglecting the structural constraints of biological laws. This reliance on purely data-driven feature extraction often yields biologically incoherent predictions and limited mechanistic interpretability. iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks. The iDCF architecture employs a dual-stream design, synergizing a standard deep network with a knowledge-based sparse neural network (KSNN) explicitly masked by pathway definitions and protein-protein interaction (PPI) networks. In comprehensive benchmarks, iDCF achieves top-tier performance, consistently ranking among state-of-the-art methods in accuracy and robustness. iDCF integrates the SHapley Additive exPlanations (SHAP) framework, bridging the gap between computational inference and biological intuition. The model's decision logic is governed by established biological mechanisms rather than spurious statistical correlations, validating its reliability. Validations across clinical contexts, including Alzheimer's disease, ovarian cancer, and diabetes, demonstrate iDCF's ability to recover disease-relevant cellular dynamics. iDCF offers a high-performance, interpretable, and biologically grounded tool for deconvolving cell-type proportions, facilitating deeper insights into tissue heterogeneity in health and disease.
A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al.· bioRxiv· 0 citations
Abstract Motivation Cell type annotation in spatial transcriptomics (ST) is fundamental for deciphering complex tissue organization and spatially resolved biological processes. Most existing methods perform ST cell type annotation by transferring labels from single-cell RNA-seq (scRNA) data to ST data, but typically rely on weakly constrained representations that neglect structured spatial dependencies and treat marker gene selection as an isolated preprocessing step. This renders them vulnerable to substantial domain gaps as well as platform-specific noise, resulting in unstable predictions and limited biological interpretability. Results To address these issues, we propose Prior-enhanced Inference for Spatial Transcriptomic Cell Type Mapping (PRISM), a novel three-stage framework integrating biological prior construction, pseudo-label generation, and multi-level ST refinement. First, PRISM constructs a cross-domain biological prior to explicitly extract marker genes to enforce positive biological discriminability. Next, it adopts a prior-enhanced self-training strategy, where scRNA-trained ensembles generate reliable pseudo-label candidates for ST data, serving as a robust anchor for cross-domain adaptation. Finally, the framework consolidates high-quality ensemble predictions selected via metric-guided evaluation, encodes spatial information, and optimizes the model under dual-directional biological constraints. Extensive experiments on eleven ST datasets across six platforms, two species, and multiple tissue contexts validate PRISM. Specifically, on the five labeled benchmarks, PRISM shows strong overall performance under both Accuracy and Macro-F1 evaluation across brain and non-brain tissues. Moreover, under fully label-free settings, PRISM achieves the best overall composite rank across all datasets, demonstrating strong robustness to domain shift and platform heterogeneity. Availability and implementation PRISM is available at https://github.com/lilab-ai4s/PRISM and https://doi.org/10.5281/zenodo.20529683.
Yiheng Xu, Xuehao Wang, Shuqi Liu et al.· Bioinformatics· 0 citations
Per-dataset analysis reveals three reproducible regimes: probabilistic variational autoencoder variants help on the smallest datasets, deep autoencoders win on mid-scale data with multi-batch or many-type structure, and classical PCA pipelines remain competitive when linear projection already captures the dominant variation.
Phong T. Nguyen, T. Vu, Thu Ha Nguyen et al.· 0 citations
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
Manoj M Wagle, Yongheng Wang, Soham Samanta et al.· bioRxiv· 0 citations
Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.
Single-cell multimodal omics offer unprecedented resolution of cellular networks, yet translating continuous computational attributions into structured, testable biological mechanisms remains a persistent bottleneck. To address this limitation, we introduce an analytical pipeline employing decision trees to discretize continuous neural network attributions into explicit regulatory thresholds. These boundaries then structurally constrain large language models, enabling them to integrate established literature with empirical data to synthesize context-specific hypotheses. Applying this continuous-to-discrete framework across sparse datasets yielded novel biological mechanisms. Specifically, the framework articulated a cytoskeletal gating hierarchy governing EGF-stimulated pathways, identified transcriptomic drivers of input resistance in cortical interneurons, and delineated translational logic predicting Ki-67 abundance within spatial transcriptomics. Retrospective benchmarking validated the framework’s capacity to autonomously reconstruct published regulatory logic. Supported by a locally deployable open-weight language model and a code-free interface, this approach establishes an auditable methodology to extract robust experimental hypotheses from high-dimensional single-cell data.
J. Chen, Yunqi Hong, Alexandra Bermudez et al.· bioRxiv· 0 citations