Skip to content
Open access

Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

Jul 2026 · bioRxiv · 0 citations
Biology Medicine

Abstract

Long noncoding RNAs (lncRNAs) play roles in gene regulation across kingdoms of life. However, lncRNAs with related functions often lack linear sequence similarity, making it difficult to leverage studies of one lncRNA to inform the understanding of others. We describe a k-mer-based hidden Markov model, hmSEEKR, that enables the scanning of transcriptomes for regions of non-linear sequence similarity to a query domain, without prior knowledge of where within the transcriptome the similarities may be located. When individual lncRNA domains were used as search features, hmSEEKR successfully identified regions in other RNAs that harbor non-linear sequence similarity and bind similar sets of proteins. Applying hmSEEKR to transcriptome-wide searches, we found that certain domains within the lncRNAs XIST, NEAT1, and MALAT1 exhibited widespread regional similarity to both lncRNA and protein-coding genes, while others were more unique, exhibiting similarity to ∼100 genes or fewer. Combinatorial searches uncovered RNAs containing sequential matches to core functional domains of XIST and NEAT1, and eCLIP-inferred protein-interaction networks within these RNAs more closely resembled those of XIST and NEAT1, respectively, than would be expected by chance, suggesting the searches recovered RNAs with similar biological properties. Finally, within annotated sets of cis-activating and cis-repressive lncRNAs, we observed opposing enrichments for similarity to domains associated with transcription-promoting complexes and heterogeneous nuclear ribonucleoprotein (hnRNP) binding, respectively, suggesting the enriched sequences may contribute to regulatory functions. hmSEEKR can be applied with minimal training data and enables the a priori discovery of RNA domains that share nonlinear similarity, offering a sequence-informed approach to discover functional elements within noncoding transcriptomes.

Read PDF

Similar papers

2026

A Comprehensive Pipeline for Long Non-Coding RNA Discovery and Characterization in Cancer.

It is proposed that the upregulated lncRNA ENSG00000265613 may enhance malignancy by stabilizing the RNA target ENSG00000582008 in luminal A breast cancer, particularly given its established role in oncogenesis.

C. Guda, Sankarasubramanian Jagadesan, Avinash M. Veerappa · 0 citations
Open access Jul 2026

An integrative workflow for lncRNA orthology detection and its application to 13 evolutionarily diverse species

Abstract Long non-coding RNAs (lncRNAs), transcripts longer than 200 nucleotides with limited protein-coding potential, are key regulators of gene expression, yet their evolutionary conservation remains poorly understood due to rapid sequence divergence. We present a flexible workflow for cross-species inference of lncRNA orthology combining two synteny-based approaches with multi-species genome alignment-derived sequence conservation. The workflow relies on standardized genome annotations and one-to-one orthologous protein-coding gene relationships, retrieved here from Ensembl resources. Applied to 13 vertebrate species spanning zebrafish, birds, and mammals, chosen to capture both broad phylogenetic distances and heterogeneous genome annotation quality, and using human (18 859 lncRNAs) as reference, the approach identified on average ∼200 putative orthologs per species under stringent criteria and up to ∼5000 under relaxed criteria. At the multi-species levels, >450 human lncRNAs were conserved in at least two species under stringent conditions, and over 10 000 in at least five species under relaxed criteria. Functional downstream analyses further revealed partial conservation of expression across 17 homologous tissues between human and chicken, as well as conserved short sequence motifs detected with LncLOOM. Together, this study provides both an adaptable workflow and a multi-species atlas to investigate lncRNA conservation and prioritize candidates for functional studies.

Fabien Degalez, Coralie Allain, L. Lagoutte et al. · 0 citations
Open access Jul 2026

Signatures of micropeptides encoded by lncRNAs in cancer progression and metastasis.

Tr-lncRNA-derived MPs represent a previously underexplored class of potentially functional molecules associated with cancer clinical annotation and may serve as biomarkers for disease progression.

Stav Zok, M. Linial · 1 citation
Open access Aug 2026

Bioinformatics analysis of regulatory relationships between long non-coding RNA genes and transcription factor coding genes

Long non-coding RNAs (lncRNAs) are increasingly recognized as important regulators of gene expression, yet their interactions with transcription factors (TFs) remain poorly understood. While previous studies have identified lncRNA-TF associations in specific biological contexts, a broader perspective on their evolutionary and functional relationships remains necessary. In this study, we systematically analyze the proximity and co-expression patterns of lncRNAs and TF-encoding genes in humans. We reveal consistent spatial associations and tissue-specific co-expression between lncRNAs and TFs using genome-wide annotations and transcriptomic data. Time-series analysis reveals interesting, but not definitive, dynamic correlations suggesting potential functional interactions. Additionally, our evolutionary analysis identified conserved pairs across species, particularly those related to developmental processes such as eye development, highlighting possible avenues for further investigation in evolutionary and developmental contexts. In summary, our findings provide new insights into the spatial and co-expression associations between lncRNAs and TF genes, and suggest directions for future research on the potential roles of lncRNAs in gene regulatory networks.

T. Yamada, Martin Loza, I. Hatada et al. · 0 citations
Open access Aug 2026

Pangenome discovery and characterization of human protein-coding duplicated genes

Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

Luyao Ren, DongAhn Yoo, Katarina Vlajic et al. · 0 citations