Skip to content

Cell Type-specific Isoform Function Prediction by Multiplex Heterogeneous Network.

Aug 2026 · IEEE transactions on computational biology and bioinformatics · Vol PP, pp. 1-12 · 0 citations
Medicine

TL;DR

Sequence-importance analysis highlights critical amino-acid regions and shows that domains annotated with the same function can exhibit distinct importance profiles across spliced isoforms, providing new insights into cell-type-specific isoform functionality and establishing cIsoFun as a practical tool for single-cell isoform analysis.

Abstract

Alternatively spliced isoforms from the same gene can perform distinct functions; however, their cell-type-specific roles remain largely uncharacterized, limiting our ability to understand cellular diversity and development beyond traditional gene-level analyses. We present cIsoFun, a multi-modal fusion framework for cell-type-specific isoform function prediction from single-cell transcriptomics data. cIsoFun leverages pre-trained ESM-2 and BERT models to extract initial sequence features, constructs a multiplex heterogeneous network over genes, isoforms, GO terms, and cell types to represent their complex relationships, and applies relation-aware attention to integrate multi-modal information and refine node embeddings. It then optimizes a multi-component loss on the updated embeddings to predict isoform functions, enabling biological interpretability via sequence-importance and cell-type-specific analyses. Experiments demonstrate that cIsoFun outperforms existing methods, particularly for sparse GO terms, and reveal distinct functional programs across contexts: kidney tumor cells are enriched for metabolism and growth regulation, skin tumor cells emphasize immune surveillance and migration, and cell lines prioritize DNA repair and telomere maintenance. Sequence-importance analysis highlights critical amino-acid regions and shows that domains annotated with the same function can exhibit distinct importance profiles across spliced isoforms. Together, these results provide new insights into cell-type-specific isoform functionality and establish cIsoFun as a practical tool for single-cell isoform analysis. Code and datasets are available at www.sdu-idea.cn/codes.php?name=cIsoFun.

View source

Similar papers

Open access Jul 2026

FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations

A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).

Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al. · 0 citations
Open access Aug 2026

ScanNet: Single-cell annotation informed by transcriptional regulation Network via iterative heterogeneous graph learning

Accurate annotation of cell types in single-cell transcriptome sequencing (scRNA-seq) data is critical for understanding cellular identities. The transcriptional regulatory networks (TRNs), which map the regulatory relationships between transcription factors (TFs) and their target genes (TGs), capture the molecular dependencies underlying transcriptional programs. However, most existing cell type annotation methods do not fully exploit this regulatory information. Therefore, we introduce ScanNet, a Single cell annotation method informed by transcriptional regulation Network, to integrate prior knowledge of TRN into data of gene expression and capture the cell-type-specific characteristics underlying TRN mechanism. TRN can be naturally represented as heterogeneous graphs consisting of two regulatory elements, TFs and TGs connected by directed edges, thereby encoding the regulatory dependencies that shape transcriptional programs and ultimately determine cellular identity. To leverage this structure, ScanNet introduces an iterative heterogeneous graph convolutional framework that learns both local and global cellular embeddings through a dual-channel encoder. The Regulation-level Encoder applies iterative heterogeneous graph convolution to capture local TF-TG regulatory interactions within TRN, while the Expression-level Encoder learns global cellular transcriptional states. By integrating the multiple-view representations, ScanNet can accurately annotate cell types. Comprehensive evaluations across eight scRNA-seq datasets spanning different species, sample scales, and sequencing platforms demonstrate that ScanNet consistently outperforms ten state-of-the-art cell type annotation methods. By embedding prior TRN structures into a heterogeneous graph, ScanNet also achieves robust performance in cross-platform cell type annotation and in identifying novel cell types under constrained structural information. Moreover, the ScanNet framework can be flexibly transferred to single-cell ATAC-seq (scATAC-seq) data by mapping chromatin accessibility to gene level, where it achieves superior performance compared to existing annotation tools. Overall, ScanNet is a scalable, transferable, and mechanistically informed framework for accurate cell type annotation across diverse single-cell data modalities.

Yongyu Long, Wenhao Zhang, Lan Cao et al. · 0 citations
Open access Jul 2026

A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence

The nervous system comprises highly diverse cell types governed by cell-type-specific molecular regulatory programs. However, regulatory evidence is scattered across unstructured literature and described using inconsistent cell-type nomenclature and granularity, hindering systematic integration and cross-study comparison. Here, we construct a neural-cell-centric multimodal knowledge graph that transforms fragmented regulatory evidence into a standardized, computable substrate. We establish a three-level hierarchical cell-type taxonomy anchored to the Cell Ontology (79 nodes), integrate two large-scale human brain single-cell transcriptomic datasets (over 4 million cells) to derive molecular fingerprints, and use a large language model to retain 25,812 curated regulatory evidence records from PubMed abstracts. The resulting Neo4j graph contains 41,532 directed relationships. For knowledge graph embedding, we export a deduplicated non-paper training subgraph containing 19,819 triples over 10,660 entities, supporting cell-type-specific link prediction that prioritizes candidate regulators and markers, illustrated here for microglia. This framework provides a structured basis for cross-study comparison, hypothesis generation and knowledge-guided reasoning in neural cell-type-specific regulation.

Chuangyu Chen, Xiaomin Ni, Yang Min et al. · 0 citations
Open access Jul 2026

CASCADE recovers promoter-associated regulatory motifs from cell-type-resolved DNA language-model attributions

Gene expression is governed by regulatory DNA and their associated trans factors acting in specific cell types, yet the sequences underlying this control remain poorly mapped in plants. Genome-pretrained DNA language models provide a route to interrogate regulatory sequence directly, but their attributions have largely been interpreted using bulk or whole-tissue data, and standard attribution pipelines can preferentially highlight sequences downstream of the transcription start (TSS) site rather than promoter-associated signals. Here, we train a celltype-resolved sequence-to-expression model from a single-cell soybean (Glycine max) atlas by coupling a soybean-adapted Genomic Pre-trained Network (GPN) to a shared sequence encoder with 66 cell-type-specific output heads. Across 38,339 protein-coding genes, the model achieves a mean per-cell-type, across-gene Pearson correlation of 0.683 and, recast as a highversus-low expression classification, reaches an area under the ROC curve of 0.92 to 0.97 across tissues, at or above dedicated plant sequence models. We then introduce ContextAware Significance of Cross-gene Attribution for Discovering Elements (CASCADE), a positionspecific statistical framework for identifying model-derived candidate regulatory elements from in silico saturation mutagenesis. Relative to the pooled null used by TF-MoDISco, CASCADE shifts motif recovery from downstream of the transcription start site toward promoter sequence, with 77% of CASCADE-exclusive motifs, compared with 12% of TF-MoDISco-exclusive motifs, falling within the promoter. Applied across the atlas, CASCADE identifies approximately 1.39 million candidate elements spanning broadly active, tissue-restricted and cell-type-restricted classes. Together, these analyses establish a position-aware approach for extracting promoterassociated regulatory hypotheses from sequence models and generate a cell-type-resolved map of candidate cis-regulatory elements.

Ali Farghadan, Robert J. Schmitz, Scott A. Jackson et al. · 0 citations
Jul 2026

Functionally Guided Graph Learning for Robust Cross-Patient Cell-Type Annotation in Single-Cell RNA Sequencing.

Experiments on three cross-patient scRNA-seq data sets demonstrate that PathoGraph achieves stable annotation performance across 32 directed reference-to-query transfer tasks, showing competitive and stable performance compared with representative marker-based, correlation-based, and model-based annotation methods.

Yue C. Li, Mengmeng Wei, Xinfei Wang et al. · 0 citations
Open access Jul 2026

Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction

Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue–cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122–0.142 for sequence-derived approaches and 0.081–0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution–pseudobulk agreement for sequence-derived models (r = 0.35–0.43 across tissue–cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.

Stephen Sim, Li Shen · 0 citations