Aug 2026· Algorithms· Vol 19, pp. 642· 0 citations· 46 references
TL;DR
This study introduces a novel sequence-based framework for PPI prediction, which combines position-specific scoring matrices (PSSMs), 3D local optimal orientation patterns (3Dloop), and histogram gradient boosting (HistGB) and shows that the approach provides a reliable and efficient solution for PPI prediction.
Abstract
Protein–protein interactions (PPIs) are fundamental to cellular processes, and understanding their mechanisms aids in disease diagnosis, drug target identification, and therapeutic development. Traditional experimental methods for PPI detection are costly and time-consuming, highlighting the need for efficient computational tools. In this study, we introduce a novel sequence-based framework for PPI prediction, which combines position-specific scoring matrices (PSSMs), 3D local optimal orientation patterns (3Dloop), and histogram gradient boosting (HistGB). Protein sequences are first transformed into PSSMs to capture evolutionary conservation, which are then processed into folded PSSMs (FPSSMs) to reveal hidden relationships among discontinuous amino acids. High-dimensional features are extracted using 3Dloop and classified with HistGB. We demonstrate the superiority of this method over random forest (RF) and support vector machine (SVM) models, achieving accuracies of 95.61% on the yeast dataset and 89.93% on the Helicobacter pylori dataset. Ablation studies confirm the effectiveness of each component in the framework. The results show that our approach provides a reliable and efficient solution for PPI prediction.
SPPIPred, an advanced machine learning-based model designed for precise PPI prediction, is presented, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.
M. Rahman, M. Ali, Md. Shohidullah et al.· PLoS ONE· 0 citations
A hybrid ensemble framework integrating XGBoost, convolutional neural networks (CNN), and graph neural networks (GNN) trained on a curated SKEMPI v2.0 dataset provides a robust and practical tool for ΔΔG prediction with potential applications in protein engineering and rational mutation design.
Sowmya Hari, R. Babu· Computational biology and ch...· 0 citations
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
The novel combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture provides a robust and biologically informed framework for PPI prediction.
Disordered proteins (IDPs) and disordered protein regions (IDRs) have
important roles in cellular signalling and regulation and in disease development. However, their
flexible conformations create challenges in annotating them through computational methods. Recent
advances in IDP disorder predictive methods based on deep learning models (e.g., SPOTDisorder,
AUCpreD, IDP-Fusion) have improved the accuracy of IDP disorder prediction; however,
all current methods still require heavy computational resources and are lacking in the ability
to interpret their results.
A new lightweight method for predicting IDRs called IDP-T5CKNN, which uses embeddings
produced by ProtT5-XL-UniRef50 and a cosine similarity k-nearest neighbour (KNN) classifier
to predict the residue-level disorder in an input protein sequence, is proposed in this study.
The residue level embeddings have been normalised and class-balanced, and evaluated using
standard binary classification metrics. This study also estimated the computational costs for the
IDP-T5CKNN method using formal Big-O notation.
For the independent MXD494 dataset, the IDP-T5CKNN method had a maximum correlation
coefficient (MCC) of 0.5459 and a balanced accuracy coefficient (BAC) of 0.7906, outperforming
all other currently available IDP disorder predictors. For the sample from the fiDPnn
Test176 dataset, the IDP-T5CKNN method produced an MCC of 0.3942 and a BAC of 0.7163, with
approximately equal sensitivity and specificity, while also achieving a comparable performance to
deep neural networks without requiring iterative training.
The results of this study demonstrate that embeddings generated by protein language
models contain disorder-relevant information and that classification methods based on similarities
to other proteins can achieve similar performance levels as deep artificial neural networks, while
also improving our ability to understand the meaning of the outputs and reducing the amount of
computational resources needed to produce the desired results.
Overall, the IDP-T5CKNN method provides a low-cost, scalable, and interpretable
method for making residue-level disorder predictions for entire proteomes.
Deepak Chaurasiya· Current Computer Science· 0 citations
It is demonstrated that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools, and shown that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions.
G. Rajagopal, Søren C. Spina, Joe Bailey et al.· bioRxiv· 0 citations