Skip to content
Open access

Comprehensive Evaluation of Protein Language Model Embeddings for Drug–Target Affinity Prediction

Sep 2026 · bioRxiv · 0 citations · 68 references
Biology

TL;DR

Results indicate simple architectural modifications to traditional convolution methods may be sufficient to bridge the gap to large pre-trained PLMs, and an architectural modification of the baseline convolution method designed to improve protein representation learning is evaluated.

Abstract

Accurate identification of drug–target interactions is consequential for novel drug discovery and development. Deep learning methods for drug–target affinity (DTA) prediction have shown great promise in accelerating drug discovery and reducing development costs. Although graph neural networks have improved drug representation learning for DTA prediction tasks, many models still struggle to effectively and efficiently capture protein information, limiting overall prediction accuracy. In this work, we systematically evaluate the impact of pre-trained protein language models (PLMs) on the downstream task of predicting binding affinity between drugs and target proteins. We design multiple experiments across four different molecular representation backbones and assess the effect of incorporating PLM embeddings, comparing their performance to classical 1D convolution methods. We evaluate four families of PLMs which we integrate into PLM-GraphDTA, each built on distinct architectures and optimized for different tasks, including structure prediction, function prediction, and sequence unmasking. Additionally, we evaluate DeepGraphDTA, an architectural modification of the baseline convolution method designed to improve protein representation learning. The models are evaluated on two benchmark datasets, Davis and KIBA, using concordance index (CI) and mean squared error (MSE) as performance metrics. We further evaluate the generalization power of each model using cold-start train and test splits, and analyze the per-protein contribution to total CI. The results indicate simple architectural modifications to traditional convolution methods may be sufficient to bridge the gap to large pre-trained PLMs.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Drug-Target Interaction Prediction via Hierarchical Sequential Cross-Attention over Chemical and Protein Language Models

Predicting Drug-Target Interactions~(DTIs) is a central task in computational drug discovery, with direct applications in virtual screening, drug repurposing, and therapeutic candidate prioritization. Although recent deep learning methods have improved DTI prediction, many sequence-based models still process drugs and...

Khadidja Henni, Hamza Abdelali, Abdelkrim Aries et al. · 0 citations
Review Aug 2026

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

It is indicated that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings, and most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization.

J. Khan, Md Hossain Shuvo · 0 citations
Aug 2026

IHLO-DTI: Drug-Target Interaction Prediction Based on Improved Hypergraph Neural Network and Laplacian Matrix Optimization

IHLO-DTI, a novel prediction model based on an improved hypergraph neural network and Laplacian matrix optimization, can effectively capture high-order many-to-many interactions between drugs and targets, improving prediction accuracy and robustness.

Guolongwei Dai, Tao Luo, Dan-Dan Li et al. · 0 citations
Open access Aug 2026

InfoMedex: drug-drug interaction prediction via a multimodal CNN-transformer model

This work investigated case studies for interactions with bupropion and ritonavir with integrated gradients and identified molecular regions associated with known CYP-mediated interaction mechanisms.

Andrew Disharoon, Shifi Pasupuleti, Clark Thurston et al. · 0 citations
Open access Sep 2026

Retrieval-Guided Transfer Learning for Low-Resource Ebola Drug–Target Affinity Prediction

Drug–target affinity (DTA) prediction plays an important role in computational drug discovery; however, its application to emerging infectious diseases such as Ebola remains challenging because of the limited availability of experimentally measured affinity data. To address this low-resource setting, we propose a retri...

Mubarakah Alotaibi, Nada Al Taweraqi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.