DeepPNI is a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes, developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes.
Abstract
Abstract Protein–nucleic acid interactions (PNIs) are central to fundamental biological processes, and mutations can disrupt these interactions by altering local structural features and binding free energy. Here, we present DeepPNI, a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes. The model was developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes, representing one of the largest curated datasets for PNI binding free energy prediction. Structural features were encoded using an edge-aware relational graph convolutional network, while sequence features were represented using the Evolutionary Scale Modeling 2 protein language model. Despite the increased dataset size and heterogeneity, DeepPNI achieved an overall Pearson correlation coefficient of 0.76 in five-fold cross-validation. Consistent performance was observed across protein–DNA and protein–RNA subsets, datasets grouped by experimental temperature, and external blind test datasets, suggesting robustness against dataset heterogeneity. DeepPNI is freely available as a web server at https://research.iitbhilai.ac.in/molinfo/deeppni.
These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.
Rozeena Arif, Alfredo Castello· Current Opinion in Structura...· 0 citations
A hybrid ensemble framework integrating XGBoost, convolutional neural networks (CNN), and graph neural networks (GNN) trained on a curated SKEMPI v2.0 dataset provides a robust and practical tool for ΔΔG prediction with potential applications in protein engineering and rational mutation design.
Sowmya Hari, R. Babu· Computational biology and ch...· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
iSCALE serves as an effective in silico tool for large-scale protein-RNA binding ΔΔG prediction, which pushes the border of understanding in mutation-induced pathological outcomes.
Protein-RNA interactions (RPIs) stand for the central process in post-transcriptional regulation and have catalyzed a fast proliferation of computational approaches in recent years. Adopting a task-oriented classification method, RPIs calculation prediction schemes proposed over the period 2010-2025 fall into five primary categories: RNA-binding protein (RBP) classification, RPIs prediction, binding site and binding profile modeling on RNA, residue-level RNA-binding interface prediction on proteins, and quantitative estimation of binding affinity and mutation effects. This study reviews the methodological evolution from conventional machine learning to deep learning, graph neural networks and large-scale pre-trained language models, and compares their differences in data preparation, evaluation protocols and generalization behavior. Particular emphasis is placed on recent advances in structure-aware and condition-aware models, as well as learning in low-data regimes. Finally, the study outlines practical recommendations for field-wide benchmarking and looks ahead to the integration with spatial omics and the development of dynamic, generative landscapes of RPIs to better empower biomedical research.