A scalable technology is introduced that quantifies thousands of proteins and transcripts in the same cell, separating RNA from protein by tip-based C18 capture and pairing full-length RNA sequencing with latest-generation mass spectrometry to yield deep proteome and transcriptome from the same cells.
Abstract
Bulk transcriptome and proteome correlate only modestly, but this has not been investigated in the same cell or across cell-state changes. Here we introduce a scalable technology that quantifies thousands of proteins and transcripts in the same cell, separating RNA from protein by tip-based C18 capture and pairing full-length RNA sequencing with latest-generation mass spectrometry. In HeLa cells, transcript and protein abundances agree on the broad ranking within a cell (r = 0.45), but do not co-vary across the population (r = 0.038). In pluripotency transitions, only a third of matched transcripts and proteins change synchronously, yet the transcription factors defining each state stay tightly co-regulated. Transcript variance is several-fold larger than protein variance, reflecting transcriptional bursting and mRNA sampling noise. The proteome is thus the stable, low-noise definition of cell state, while the transcriptome marks cellular transitions; consequently, the proteome defines cell-state from far fewer cells. Highlights Multimodal workflow yields deep proteome and transcriptome from the same cells Transcript and protein rank similarly within a cell but are uncoupled across cells Only a third of RNA-protein pairs change together, yet state-defining TFs stay coupled Protein varies several-fold less than mRNA, defining state from far fewer cells
Differential expression studies, single-cell atlases, and target discovery pipelines measure RNA as a proxy for protein. Across human tissues, RNA explains only 40 to 60% of the variance in protein, and no framework resolves which genes reliably translate RNA into protein. We present an integrated atlas that matches single-cell transcriptomics from Tabula Sapiens and the Human Brain Cell Atlas to cell type-resolved immunohistochemistry from the Human Protein Atlas, comprising 488,190 observations across 11,154 genes, 24 tissues, and 53 cell types. For each gene we defined a suppression rate and separated genes into concordant, variable, and suppressed classes. The pooled correlation of ρ ≈ 0.4 reflects the mixing of these classes, and concordance depends on the interaction between gene and tissue. Suppression is predictable from gene sequence alone and traces to reduced translational efficiency and assembly-dependent degradation rather than to mRNA decay. The suppressed class is enriched for drug targets nominated on clinical evidence but not those validated by compound activity. We implement these classifications in an R package, concordR, and audit proteins nominated as brain-derived targets in neurodegeneration, where almost none survive at the protein level. Our atlas establishes RNA to protein concordance as a measurable property of the individual gene.
Abstract Transcription—the process by which genomic DNA is converted into RNA—is a highly dynamic and tightly controlled process across all domains of life. Although bacteria were once regarded as relatively simple organisms, their transcriptomes are now recognized to be remarkably complex, heterogeneous, and subject to multilayered regulation. Despite the availability of an abundance of sequenced bacterial genomes, a comprehensive understanding of how bacteria tune their transcriptional output to adapt to changing environments remains lacking. To this end, SEnd-seq (simultaneous 5′ and 3′ end sequencing) was developed as a high-throughput approach uniquely capable of simultaneously capturing both 5′ and 3′ ends of individual RNA molecules, enabling the reconstruction of full-length transcripts. By capturing each RNA molecule as a distinct molecular entity with single-nucleotide resolution, SEnd-seq has uncovered previously unrecognized transcriptional features across diverse bacterial species, including even the well-studied Escherichia coli. This method performs robustly across a wide range of RNA species and organisms, including hard-to-lyse pathogens such as Mycobacterium tuberculosis. Moreover, SEnd-seq exhibits high sensitivity for detecting low-abundance RNA and is compatible with various target RNA enrichment strategies, as well as genetic, chemical, and functional perturbations, enabling context-specific transcriptomic analyses. As a versatile and broadly adaptable technology, SEnd-seq provides comprehensive insights into transcriptional regulation, RNA processing, and the coordination between RNA-based processes, thereby uncovering potential targets for antibiotic development. In the present review, we summarize the features and applications of SEnd-seq and discuss its future methodological development and expansion into broader biological and biomedical research contexts.
Single-cell RNA sequencing technology dramatically changed the way we investigate transcriptomes. However, the amount and complexity of data generated by such methods poses new challenges for biologists who are trying to extract detailed insights into the genetic programs that drive cellular functions and differentiation. To provide a more intuitive understanding of cell specific gene expression programs, we developed a novel approach for exploiting scRNA-seq data that detects individual gene expression levels in each cell, by avoiding dimensional reduction methods. This was achieved by focusing our analysis on individual cells with a high sequencing coverage (above 15000 Unique Molecular Identifiers (UMIs)). Such High Coverage Cells (HCC), were found in all five C. elegans scRNA-seq datasets we investigated and constitute direct quantitative experimental observations of the mRNA content of individual cells. Clustering the complete gene expression matrix for these cells, we identified gene sets specific for most C. elegans tissues. Among each set we found genes that are dominating cell specific transcriptomes as well as genes that are restricted to particular cell types but are a thousand fold less expressed. For each cell type or subtype we characterized, we identified a set of genes with expression restricted to those cells that were not previously associated with the corresponding tissue. Our results demonstrate that by focusing on HCCs, we can provide high-resolution quantitative descriptions of cellular expression landscapes that are immediately exploitable for researchers to generate new biological hypotheses. Overall, we demonstrate that HCCs represent a powerful and largely unexplored source of biological insights and suggest that future scRNA-seq experiments could benefit from focusing on HCC enrichment to capture and exploit the full complexity of cellular transcriptomes.
Florian Bernard, Emma Kandel, D. Dargère et al.· bioRxiv· 0 citations
This review examines the experimental and computational foundations of scLR-seq, including platform selection, library design, cell barcode and unique molecular identifier recovery, transcript discovery, and isoform quantification, and summarize emerging insights into isoform usage, alternative splicing, transcription start and end site selection, allele-specific expression, fusion transcripts, transposable element-derived transcripts, and RNA modifications.
Cell and tissue functions arise from complex interactions among numerous genes, and a systematic understanding of these functions requires isoform-resolved transcriptomic analysis of single cells with high spatial resolution. Here, we introduce an in situ RNA amplification method and its integration with multiplexed error-robust fluorescence in situ hybridization (MERFISH) to detect short RNA sequences and enable whole-transcriptome-scale, isoform-resolved spatial transcriptomics of individual cells in intact tissues. Using this approach, we imaged ∼33,000 distinct RNAs-including ∼23,000 genes and ∼10,000 isoforms-in the mouse brain. Our data enabled systematic analyses of region- and cell-type-specific gene programs and ligand-receptor-based cell-cell communications. These data further revealed rich spatial diversity and cell-type specificity in isoform usage across numerous genes, as well as brain structures particularly rich in isoform specificity. We anticipate broad application of this method for characterizing the molecular and cellular basis of tissue functions, unlocking previously inaccessible discoveries in cell and organismal biology.
Limor Cohen, Aaron R Halpern, Timothy R. Blosser et al.· Cell· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.