Aug 2026· Current Opinion in Structural Biology· Vol 101, pp.
103362
· 0 citations· 55 references
Medicine
TL;DR
These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.
Abstract
RNA-binding proteins (RBPs) are essential across biology, from viruses to complex multicellular organisms. They regulate gene expression and cellular responses, making RNA recognition central to understanding health and disease. Biochemical, biophysical, and structural studies have defined core principles of RNA binding, but recent RNA interactome surveys have expanded the RBP repertoire and revealed many noncanonical RNA-binding regions. This diversity demands highly scalable predictive methods. Here, we review machine learning predictors built on protein language models and structure-aware representations. These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.
RNA structure is central to the function of every RNA class yet the gap between annotated sequences and experimentally determined structures remains large. Computational methods to fill this gap have evolved from thermodynamic free energy minimization through supervised deep learning to self-supervised RNA language models trained on millions of sequences, progressively improving structure prediction. Here we review the state of the art in RNA structure prediction, covering key training datasets, community benchmarks, and the performance of current models. We further discuss perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models. Challenges in generalization, handling of noncanonical interactions, and contextual structure prediction remain open frontiers for the field.
Lambert Moyon, Annalisa Marsico· Current Opinion in Structura...· 0 citations
Protein-RNA interactions (RPIs) stand for the central process in post-transcriptional regulation and have catalyzed a fast proliferation of computational approaches in recent years. Adopting a task-oriented classification method, RPIs calculation prediction schemes proposed over the period 2010-2025 fall into five primary categories: RNA-binding protein (RBP) classification, RPIs prediction, binding site and binding profile modeling on RNA, residue-level RNA-binding interface prediction on proteins, and quantitative estimation of binding affinity and mutation effects. This study reviews the methodological evolution from conventional machine learning to deep learning, graph neural networks and large-scale pre-trained language models, and compares their differences in data preparation, evaluation protocols and generalization behavior. Particular emphasis is placed on recent advances in structure-aware and condition-aware models, as well as learning in low-data regimes. Finally, the study outlines practical recommendations for field-wide benchmarking and looks ahead to the integration with spatial omics and the development of dynamic, generative landscapes of RPIs to better empower biomedical research.
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
DeepPNI is a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes, developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes.
A long-context protein language model is introduced, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors.
T. Dejean, Barbra D. Ferrell, Zachary D. Schreiber et al.· GigaScience· 0 citations
Protein-RNA complexes drive fundamental cellular processes such as transcription and translation. Despite the prevalence and importance of protein-RNA interactions, the field lacks reliable and accessible methods to quantify the energetic favorability of these interactions. We propose an experimentally tuned protein-RNA score function that can be directly implemented into ROSETTA. Fine-tuning these score functions for predictive tasks requires repeated evaluations on a set of protein-RNA complexes, which can be computationally expensive given the number of parameters to tune. We used Bayesian Optimization to efficiently improve the energetic agreement between ROSETTA and experimentation. We observe significant interactions for specific RNA subclasses, serving as further confirmation of the physical validity of the score function. Beyond protein-RNA interaction prediction, we establish a framework to efficiently fine-tune ROSETTA score functions for any protein-class interaction using Bayesian Optimization. TOC FIGURE
Joe Bailey, Nathan Phan, Søren C. Spina et al.· bioRxiv· 0 citations