Skip to content
Open access

DeepPNI: a language- and graph-based model for mutation-driven protein–nucleic acid binding energetics

Aug 2026 · Nucleic Acids Research · Vol 54 · 0 citations · 50 references
Medicine

TL;DR

DeepPNI is a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes, developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes.

Abstract

Abstract Protein–nucleic acid interactions (PNIs) are central to fundamental biological processes, and mutations can disrupt these interactions by altering local structural features and binding free energy. Here, we present DeepPNI, a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes. The model was developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes, representing one of the largest curated datasets for PNI binding free energy prediction. Structural features were encoded using an edge-aware relational graph convolutional network, while sequence features were represented using the Evolutionary Scale Modeling 2 protein language model. Despite the increased dataset size and heterogeneity, DeepPNI achieved an overall Pearson correlation coefficient of 0.76 in five-fold cross-validation. Consistent performance was observed across protein–DNA and protein–RNA subsets, datasets grouped by experimental temperature, and external blind test datasets, suggesting robustness against dataset heterogeneity. DeepPNI is freely available as a web server at https://research.iitbhilai.ac.in/molinfo/deeppni.

Read PDF

Similar papers

Review Open access Aug 2026

A new dimension in protein-RNA interface prediction: Integrating protein language models and geometric deep learning.

These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.

Rozeena Arif, Alfredo Castello · 0 citations
Aug 2026

A hybrid CNN-GNN-XGB ensemble framework for prediction of mutation-induced protein-protein binding free-energy changes (ΔΔG) in protein engineering.

A hybrid ensemble framework integrating XGBoost, convolutional neural networks (CNN), and graph neural networks (GNN) trained on a curated SKEMPI v2.0 dataset provides a robust and practical tool for ΔΔG prediction with potential applications in protein engineering and rational mutation design.

Sowmya Hari, R. Babu · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

DHST: A Deep Hybrid Structure–Topology Framework for Accurate Protein Function Prediction

DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.

Bin Lu, Fujun Xiang, Hai-Long Wang et al. · 0 citations
Review Jul 2026

From binary labels to dynamic landscapes: The evolving computational prediction of protein-RNA interactions through tasks and deep learning paradigms.

Protein-RNA interactions (RPIs) stand for the central process in post-transcriptional regulation and have catalyzed a fast proliferation of computational approaches in recent years. Adopting a task-oriented classification method, RPIs calculation prediction schemes proposed over the period 2010-2025 fall into five primary categories: RNA-binding protein (RBP) classification, RPIs prediction, binding site and binding profile modeling on RNA, residue-level RNA-binding interface prediction on proteins, and quantitative estimation of binding affinity and mutation effects. This study reviews the methodological evolution from conventional machine learning to deep learning, graph neural networks and large-scale pre-trained language models, and compares their differences in data preparation, evaluation protocols and generalization behavior. Particular emphasis is placed on recent advances in structure-aware and condition-aware models, as well as learning in low-data regimes. Finally, the study outlines practical recommendations for field-wide benchmarking and looks ahead to the integration with spatial omics and the development of dynamic, generative landscapes of RPIs to better empower biomedical research.

Xinyu Li, Qianmao Wen, Zilong Zhang et al. · 0 citations