Prediction of Drug Target Affinity (DTA) is essential for accelerating computational drug discovery and reducing experimental costs. However, traditional experimental approaches for DTA estimation are resource-intensive and are further challenged by the structural flexibility of both drugs and target proteins. In this work, we propose the PCBERT-GAT-DFFNN-DTA model, a three-stage deep cross-modal representation fusion framework for accurate DTA prediction. In the first stage, variable-length protein sequences are transformed into contextual representations using ProtBERT to obtain fixed-size protein embeddings. Drug molecules are represented using two modalities: sequence-based embeddings generated from ChemBERT and structure-based embeddings learned from molecular graphs using a Graph Attention Network (GAT). In the second stage, each modality is processed through dedicated subnetworks to refine features and reduce dimensionality while preserving modality-specific information. In the final stage, the refined representations are fused and passed to a Deep Feed-Forward Neural Network (DFFNN) to predict drug target binding affinity. The proposed model consistently outperformed most baseline methods under the S1-S3 evaluation settings across the benchmark datasets. Under the more challenging S4 blind setting, the model achieved strong performance on the KIBA dataset and competitive results on the Davis and Metz datasets. Compared with LLMDTA, the proposed approach achieves significant improvements in R2 scores across all datasets, demonstrating its effectiveness in learning complex drug protein interactions for reliable DTA prediction.
Essmily Simon, Sanjay S. Bankapur· Analytical Biochemistry· 0 citations
A novel deep learning framework is proposed that leverages pre‐trained BERT‐based language models to extract contextual embeddings from protein and drug sequences and refined using a proposed dedicated ResNet‐based subnetwork to preserve intrinsic biochemical characteristics.
Essmily Simon, Sanjay S. Bankapur· Chemical Biology and Drug De...· 0 citations
Evaluating the effectiveness of embeddings from five Protein Language Models, including ProtBERT-BFD, ESM-2, ProtALBERT, ProLLaMA, and ProtGPT-2, as input features for various machine learning classifiers suggests that while current embeddings offer strong performance, further advancements in feature extraction and model architectures are needed to significantly boost strict accuracy.
Karthik Avinash, S. Tejas, Sriram Mamidala et al.· Analytical Biochemistry· 0 citations