Jul 2026· IEEE journal of biomedical and health informatics· Vol PP· 0 citations
Medicine
TL;DR
A multi-modal deep learning framework to predict drug-target affinity by integrating sequence semantics with graph structural information and design a new symmetric dual cross-attention fusion mechanism for drugs and targets.
Abstract
Accurate prediction of drug-target affinities (DTA) is critical for drug discovery. However, this task remains a significant challenge due to the complexity of modeling interactions between small ligands and large targets. In this study, we propose a multi-modal deep learning framework (CrossSG-DTA) to predict drug-target affinity by integrating sequence semantics with graph structural information. We leverage ChemBERTa and ESM-2 to extract rich semantic features for drugs and targets, respectively. In addition, a modified Graph Convolutional Network (GCN) is utilized to simultaneously capture structural data. To effectively fuse these heterogeneous features, we design a new symmetric dual cross-attention fusion mechanism for drugs and targets. This mechanism enables the model to capture complex dependencies between global sequence representations and local topological structures. Subsequently, the fused drug and target features are concatenated and fed into a three-layer Multi-Layer Perceptron (MLP) to obtain the final binding affinity. Experimental results on the Davis and KIBA datasets demonstrate that CrossSG-DTA significantly outperforms state-of-the-art methods. Finally, a case study on a glaucoma-related target highlights the practical utility of our model as a powerful in silico tool for DTA tasks.
This work proposes GraphTransDTI, a synergistic hybrid framework that integrates a Graph Transformer to represent drug graph structures, a CNN-BiLSTM network to encode protein sequence context, and a Cross-Attention mechanism to model cross-domain interactions.
Vang V. Le, Mai Thi Anh Nhu, Pham Truong Viet Thong· PLoS ONE· 0 citations
Experiments show that GraESM-FuseDTA achieves competitive overall performance and consistent advantages in ranking-oriented and variance-explanation metrics across warm start, drug cold start, target cold start, and strict pair cold start settings.
DeepGCL is presented, a novel multi-modal framework that leverages multi-view graph contrastive learning to capture latent representations of pocket-drug interactions and their underlying molecular determinants and underscores the effectiveness of multi-view learning paradigms in capturing the multifaceted nature of drug-target interactions.
Hongmei Wang, Shisen Sun, Mujin Li et al.· IEEE journal of biomedical a...· 0 citations
Prediction of Drug Target Affinity (DTA) is essential for accelerating computational drug discovery and reducing experimental costs. However, traditional experimental approaches for DTA estimation are resource-intensive and are further challenged by the structural flexibility of both drugs and target proteins. In this work, we propose the PCBERT-GAT-DFFNN-DTA model, a three-stage deep cross-modal representation fusion framework for accurate DTA prediction. In the first stage, variable-length protein sequences are transformed into contextual representations using ProtBERT to obtain fixed-size protein embeddings. Drug molecules are represented using two modalities: sequence-based embeddings generated from ChemBERT and structure-based embeddings learned from molecular graphs using a Graph Attention Network (GAT). In the second stage, each modality is processed through dedicated subnetworks to refine features and reduce dimensionality while preserving modality-specific information. In the final stage, the refined representations are fused and passed to a Deep Feed-Forward Neural Network (DFFNN) to predict drug target binding affinity. The proposed model consistently outperformed most baseline methods under the S1-S3 evaluation settings across the benchmark datasets. Under the more challenging S4 blind setting, the model achieved strong performance on the KIBA dataset and competitive results on the Davis and Metz datasets. Compared with LLMDTA, the proposed approach achieves significant improvements in R2 scores across all datasets, demonstrating its effectiveness in learning complex drug protein interactions for reliable DTA prediction.
Essmily Simon, Sanjay S. Bankapur· Analytical Biochemistry· 0 citations
This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.
Jiaqi Deng, Senyu Tang, Jijun Tang et al.· Journal of Chemical Informat...· 0 citations
ColdstartMHDTI is proposed, a two-stage framework for heterogeneous-graph-based DTI prediction that integrates sequence-derived structural representations with local and global relational information and supports candidate prioritization for downstream screening and evidence-guided hypothesis generation.
Hongyang Yang, Xiucai Ye, Huipu Han et al.· Frontiers in Chemistry· 0 citations