Skip to content
Open access

DB-IRES: a deep learning model based on ensemble learning for predicting internal ribosome entry sites.

Jul 2026 · BMC Bioinformatics · 0 citations
Medicine

TL;DR

DB-IRES serves as a reliable and precise computational tool for predicting IRES elements and enables more in-depth functional investigations of IRES biology and supports wider applications in RNA research and the development of associated therapeutics.

Abstract

Background

Internal ribosome entry sites (IRES) are cap-independent translation initiation elements present in specific viral and cellular mRNAs. They facilitate direct ribosome recruitment for protein synthesis, bypassing the need for the canonical 5' cap structure. Due to their pivotal roles in viral pathogenesis and cellular translational regulation, precise identification of IRES is crucial for advancing mechanistic studies and exploring potential therapeutic interventions. Nonetheless, manual identification is both labor-intensive and costly, while current computational methods exhibit limitations in accuracy and robustness.

Results

DB-IRES is a novel deep learning model designed for this purpose, incorporating densely connected 1D convolutional neural network blocks, bi-directional gated recurrent units, and a self-attention mechanism within an ensemble learning framework. Trained with a five-fold cross-validation strategy, the final integrated model exhibits superior and more robust predictive performance compared to existing methods on an independent test set. The model consistently achieves enhanced discriminatory capabilities across multiple evaluation metrics, thereby validating the effectiveness of its hybrid architecture and ensemble design.

Conclusions

DB-IRES serves as a reliable and precise computational tool for predicting IRES elements. Its improved performance enables more in-depth functional investigations of IRES biology and supports wider applications in RNA research and the development of associated therapeutics.

Read PDF

Similar papers

Open access Aug 2026

DNCLA: A Deep Learning Model for TFBS Identification Based on Structural and Conformational Properties of Nucleotides and Dinucleotides

Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.

Jingjue Wei, Jie Feng · 0 citations
Review Jul 2026

From binary labels to dynamic landscapes: The evolving computational prediction of protein-RNA interactions through tasks and deep learning paradigms.

Protein-RNA interactions (RPIs) stand for the central process in post-transcriptional regulation and have catalyzed a fast proliferation of computational approaches in recent years. Adopting a task-oriented classification method, RPIs calculation prediction schemes proposed over the period 2010-2025 fall into five primary categories: RNA-binding protein (RBP) classification, RPIs prediction, binding site and binding profile modeling on RNA, residue-level RNA-binding interface prediction on proteins, and quantitative estimation of binding affinity and mutation effects. This study reviews the methodological evolution from conventional machine learning to deep learning, graph neural networks and large-scale pre-trained language models, and compares their differences in data preparation, evaluation protocols and generalization behavior. Particular emphasis is placed on recent advances in structure-aware and condition-aware models, as well as learning in low-data regimes. Finally, the study outlines practical recommendations for field-wide benchmarking and looks ahead to the integration with spatial omics and the development of dynamic, generative landscapes of RPIs to better empower biomedical research.

Xinyu Li, Qianmao Wen, Zilong Zhang et al. · 0 citations
Open access Aug 2026

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a “promoter” or “non-promoter,” which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model’s ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Pi-Jing Wei, Wenkang Zheng, Yijun Gu et al. · 0 citations
Aug 2026

Deep3MVPF: Multiview Deep Framework for the Prediction of Stability and m6A in mRNA 3'UTR.

Accurate prediction of mRNA stability and identification of N6-methyladenosine (m6A) sites are central to understanding post-transcriptional regulation. Because the 3' untranslated region (3'UTR) contains both stability-associated cis-elements and many m6A sites, it provides a suitable context for modeling RNA regulatory effects. However, most existing methods rely primarily on linear sequence information and do not adequately capture higher-order topology or RNA structural context. Here, we present Deep3MVPF, a multiview deep learning framework for 3'UTR stability prediction and m6A site identification. Deep3MVPF integrates a multiscale convolutional neural network, a k-mer de Bruijn graph neural network, and a secondary-structure graph neural network to jointly model sequence, topological, and structural representations. For 3'UTR stability prediction, the model was trained and evaluated on a zebrafish (Danio rerio) mRNA degradation data set and achieved an MSE of 0.0049. For m6A site identification, it was evaluated on nine human cell line data sets and achieved an average AUC of 0.970. Attribution analysis further showed that Deep3MVPF recovered regulatory features consistent with known biology, including the destabilizing GCACUU motif and stabilizing G-rich/G-quadruplex-associated signals. These results demonstrate that integrating heterogeneous RNA representations can improve predictive modeling and facilitate interpretation of post-transcriptional regulatory grammar.

Junyi Liu, Qi Zhang, Jiangning Song et al. · 0 citations
Preprint Aug 2026

A Conditional Structure-Aware Generative Transformer for Multi-Objective Design of m1{\Psi}-Modified RNA 5'UTRs

The 5'untranslated region is a major determinant of translation initiation, and its effect becomes especially important in modified mRNA sequences, where start-codon context, cap-proximal secondary structure, upstream AUGs and upstream open reading frames, and nucleotide chemistry can alter ribosome scanning and initiation recruitment, scanning, and decoding in sequence-dependent ways. Recent computational studies have moved the field from prediction toward design, including massively trained predictive models such as Smart5UTR for m1$\Psi$-modified mRNA, broader 5'UTR generation and optimization frameworks such as UTRGAN and UTailoR, and structure-guided RNA design systems such as RhoDesign. Here, we describe a conditional generative framework for 50-nt modified-RNA 5'UTR design that optionally conditions on ribosome load, GC content, minimum free energy, and target secondary structure. The implementation uses a Transformer-based generator followed by sequence ranking and local refinement with a Smart5UTR-derived ribosome-load oracle and ViennaRNA-based folding metrics, including support for modified-base folding parameters. Across multiple simulation scenarios and experimental settings, different combinations of RL, GC, MFE, and structural constraints produced distinct performance tradeoffs, enabling ablation-based identification of the best-performing formulation.

Narges Zarnaghinaghsh, Ahmadreza Mofayezi, Byung-Jun Yoon · 0 citations
Jul 2026

Identifying RNA ac4C Modification Sites via Pseudo-Nucleotide Fingerprint Encoding and Multi-Scale Feature Integration.

DFM-ac4C, a novel computational framework designed for the accurate prediction of ac4C modification sites, significantly outperforms existing models, achieving outstanding predictive metrics, and underscores DFM-ac4C's effectiveness as a robust and efficient tool for RNA ac4C site identification.

Yiming Wang, Fan Mo, Yun Sha et al. · 0 citations