GRASSP provides a competitive framework for integrating pretrained RNA representations with spatial structural context while reducing reliance on additional handcrafted structural annotations, and is demonstrated to outperform state-of-the-art baselines.
Abstract
MOTIVATION
RNA-small molecule binding site prediction is crucial for targeted drug discovery. Sequence-based methods are efficient but often fail to capture structural dependencies between nucleotides, whereas structure-aware graph models can better represent spatial interactions but typically rely on complex structural annotations and multi-stage preprocessing pipelines. We therefore developed GRASSP, a streamlined hybrid deep learning framework that integrates pretrained RNA language model (LM) representations with adaptive graph refinement.
RESULTS
GRASSP leverages nucleotide embeddings and predicted secondary-structure features from a pretrained RNA LM to construct spatial RNA graphs, followed by a lightweight two-step graph attention refinement module with adaptive gating to capture local and contextual nucleotide dependencies. Across four benchmark datasets (TE18, HARIBOSS, TL12, and JL10), GRASSP generally outperformed state-of-the-art baselines, with improvements of up to 24.1% in AUC and 44.5% in MCC. Ablation analyses showed that pretrained RNA representations provided the dominant predictive contribution, while spatial graph refinement offered complementary but dataset-dependent benefits. These results demonstrate that GRASSP provides a competitive framework for integrating pretrained RNA representations with spatial structural context while reducing reliance on additional handcrafted structural annotations.
AVAILABILITY
Code and datasets are publicly available at https://github.com/langiocn/GRASSP, with an archival snapshot available on Zenodo at https://doi.org/10.5281/zenodo.21888291.
SUPPLEMENTARY INFORMATION
Supplementary data are available at Bioinformatics online.
RNA-binding proteins (RBPs) orchestrate a complex combinatorial regulatory “code” that governs RNA splicing, stability, localization, and translation. Learning the relationship between RNA sequences and these processes is a central challenge in genomics. Foundation models, notably RNA language models, have emerged as the dominant approach, learning general-purpose representations from unlabeled sequence at scale. While RNA language models have demonstrated impressive performance across a broad range of downstream tasks, they generally learn from sequence reconstruction objectives alone, lacking direct connections to the regulatory principles that govern RNA function. Here we introduce Parnet, an RNA foundation model trained directly and exclusively on experimental CLIP-seq data. Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence. This CLIP-seq pretraining strategy departs fundamentally from the masked-language-modeling paradigm, anchoring learned RNA representations directly in measured protein–RNA interactions rather than sequence statistics. Parnet substantially outperforms its single-task predecessor RBPNet in binding profile and motif recovery, generalizes to unseen cell types and iCLIP data, and recapitulates position-dependent splicing regulation. Frozen Parnet embeddings, without task-specific fine-tuning, match or exceed the performance of both task-specific tools, as well as larger self-supervised RNA and genomic language models across diverse downstream tasks, including RNA biotype classification, lncRNA chromatin localization, translational efficiency, splice-site recognition, intron retention, and non-coding variant effect prediction. Importantly, Parnet remains mechanistically interpretable, tracing predictions back to the specific RBPs and motifs that drive them. These results establish the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.
A modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hint with confidence, and a GNN extracts an instance-specific explanatory subgraph via a necessity-based edge-drop intervention.
K. Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras· 1 citation
The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.
Junkai Wang, G. Luo, Yun-Song Yang et al.· Bioinformatics· 0 citations
Experiments show that GraESM-FuseDTA achieves competitive overall performance and consistent advantages in ranking-oriented and variance-explanation metrics across warm start, drug cold start, target cold start, and strict pair cold start settings.
Edge Generation-guided Relation-aware Learning (EGRL) is proposed, a novel framework with several key components: implicit meta-path learning to capture relational semantics without handcrafted paths; a multi-relation-aware attention mechanism for adaptive fusion of interaction patterns; a graph generator that predicts potential ("soft") edges to support cold-start nodes; and a multi-feature fusion predictor for final interaction scoring.
Danyu Li, Ling Zhou, Rubing Huang et al.· 0 citations
Direct prior-guided graph learning can improve the robustness and biological interpretability of GRN inference in data-limited settings and demonstrate that directed prior-guided graph learning can improve the robustness and biological interpretability of GRN inference in data-limited settings.
N. Alkhateeb, Mamoun A. Awad· Frontiers in Bioinformatics· 1 citation