Skip to content
Open access

FerrGAT: Multi-Task Pre-Training of Graph Attention Networks for Low-Data Molecular Activity Prediction

Aug 2026 · Mathematics · Vol 14, pp. 3003 · 0 citations · 36 references

TL;DR

FerrGAT, a graph attention network framework that addresses molecular bioactivity in low-data regimes through domain-relevant multi-task pre-training and dual-channel molecular representation learning, provides an effective and interpretable framework for molecular property prediction in data-limited scenarios.

Abstract

Predicting molecular bioactivity in low-data regimes remains a central challenge in computational drug discovery, where labeled compounds for specialized tasks are scarce while related datasets are abundant. Here, we propose FerrGAT, a graph attention network framework that addresses this challenge through domain-relevant multi-task pre-training and dual-channel molecular representation learning. FerrGAT first pre-trains a shared GAT encoder on three mechanistically related tasks—antioxidant activity (GST inhibition, 245 compounds from ChEMBL target CHEMBL2095173; 83 active/162 inactive), kinase inhibition (3000 compounds spanning AXL, EGFR and VEGFR2 from ChEMBL; 2240 active/760 inactive), and cellular toxicity (7265 compounds from the Tox21 NR-AhR endpoint; 309 active/6956 inactive)—then transfers the learned representations to a target task via differential learning rate fine-tuning. The architecture fuses atom-level graph features from multi-head attention message passing with global physicochemical descriptors through a learned projection and provides built-in interpretability via attention weight visualization at the atomic level. We evaluated FerrGAT on ferroptosis inhibitor prediction as a representative low-data molecular classification task (1052 compounds from ChEMBL targets GPX4 and HMOX1 combined with 63 literature- and FerrDb-curated ferroptosis modulators; 409 active/643 inactive). In 5-fold cross-validation, FerrGAT achieved an AUC of 0.906, outperforming Morgan fingerprint baselines, including Random Forest (0.877), XGBoost (0.871), and an SVM (0.873). Ablation studies confirmed that domain-relevant pre-training improved AUC by 2.5% over training from scratch, while pre-training on unrelated tasks degraded performance, highlighting the importance of task-domain alignment. Applied to virtual screening of 30 FDA-approved tyrosine kinase inhibitors, the model identified Bemcentinib (AXL inhibitor, score = 0.918) as a top candidate, validated by AutoDock Vina molecular docking (−8.51 kcal/mol) and independent experimental evidence, including lipid peroxidation assays, Western blot, and cellular thermal shift analysis. These results demonstrate that domain-aware transfer learning with graph attention networks provides an effective and interpretable framework for molecular property prediction in data-limited scenarios.

Read PDF

Similar papers

Jul 2026

Meta-learning GNN with MD-informed attention for cross-species prediction of phosphoinositide-dependent kinase-1 (PdK1) inhibitors in termite control

This work bridges the gap between data-driven and physics-based approaches, providing a scalable solution for pesticide discovery when target-specific data are limited, and incorporating meta-learning and MD insights improves cross-species transferability and yields interpretable attention patterns based on biophysical principles.

Haroon · 0 citations
Open access Aug 2026

GAMT-GINE: A Graph Isomorphism Network Integrating Continuous Spatial Awareness and Multi-Task Learning for Protein–Ligand Binding Affinity Prediction

Comprehensive evaluations indicate that GAMT-GINE can effectively utilize continuous spatial information and heterogeneous affinity labels, achieving good predictive accuracy and cross-dataset generalization capability.

Jiarui Li, Hongquan Li, Di Wu et al. · 0 citations
Open access Jul 2026

SM-GAT: a safety-aware multi-task graph attention network for multi-target anti-diabetic lead discovery from natural products

Traditional Chinese Medicine constitutes a chemically diverse and pharmacologically rich reservoir of bioactive compounds, often exhibiting multi-target pharmacological properties that are potentially valuable for complex metabolic disorders such as type 2 diabetes mellitus (T2DM). However, systematic prioritization of active and safe constituents remains challenging, as therapeutic efficacy must be optimized concurrently with toxicity risk. Here, we present a safety-aware Multi-Task Graph Attention Network (SM-GAT) framework that jointly models anti-diabetic efficacy and toxicity liabilities of TCM-derived compounds within a unified architecture. Four tasks are simultaneously optimized: inhibition of dipeptidyl peptidase-4 (DPP4), inhibition of α-glucosidase, acute oral toxicity, and clinically relevant toxicity. By enabling shared molecular representation learning across heterogeneous yet biologically related endpoints, SM-GAT facilitates knowledge transfer between efficacy and safety domains. Across all prediction tasks, SM-GAT achieved competitive or superior performance compared with single-task graph neural networks and other baseline models, achieving ROC-AUC values up to 0.892 for α-glucosidase inhibition. Notably, multi-task learning yields pronounced improvements in data-limited settings, highlighting effective cross-task regularization. Large-scale virtual screening of the TCMBank library demonstrates practical applicability, enabling efficient prioritization of structurally diverse candidates with favorable predicted efficacy–safety balance. Several structurally diverse lead compounds, including ellagic acid derivatives, are identified with favorable predicted efficacy–safety balance. Furthermore, atom-level attention analysis highlighted chemically interpretable substructures associated with predicted efficacy and toxicity-related molecular representations. Collectively, this study establishes an interpretable multi-objective framework for safety-aware lead discovery, providing a computational framework for integrating traditional botanical resources into anti-diabetic lead discovery.

Jianxin Zhang, Hongyi Liu, Shengnan Guo · 0 citations
Open access Aug 2026

A hybrid high activity aware framework integrating graph attention network and transformer for half maximal inhibitory concentration prediction

Tyrosine kinase inhibitors targeting the c-KIT receptor are pivotal in the targeted therapy of malignancies such as gastrointestinal stromal tumors (GIST). The bioactivity of these inhibitors is typically quantified by the half-maximal inhibitory concentration (IC 50 ), making its accurate prediction a critical computational task for accelerating the discovery and optimization of anticancer lead compounds. To address this need, we propose HGATT—a hybrid high activity aware framework integrating Graph Attention Network (GAT) and Transformer—for high-accuracy half maximal inhibitory concentration (IC 50 ) prediction of c-KIT inhibitors. The model was trained on inhibitor data targeting c-KIT and related kinase families sourced from the BindingDB and ChEMBL databases. By extracting atom-level graph features, Morgan fingerprints, and physicochemical descriptors from SMILES strings, HGATT constructs a unified molecular representation that integrates both local structural and global information. Its architecture employs multidimensional graph attention mechanisms and gated residual modules to simultaneously capture atomic-level local structural features and macroscopic molecular properties. Stabilized training strategies, including gradient clipping, were adopted to enhance training efficiency and model robustness. On an independent test set, HGATT achieved a mean squared error (MSE) of 0.28 and a coefficient of determination (R²) of 0.57, corresponding to an approximate 44% reduction in MSE compared to the second-best baseline. Experimental results demonstrate that HGATT outperforms not only individual graph neural network (GNN)- and machine learning-based models but also other related drug-target prediction methods and baseline regression approaches, exhibiting superior predictive accuracy.

D. Ban, Lu Pan, Xueli Zhang et al. · 0 citations
Aug 2026

GraESM-FuseDTA: adaptive gated multimodal fusion of graph neural networks and protein language models for robust drug-target affinity prediction.

Experiments show that GraESM-FuseDTA achieves competitive overall performance and consistent advantages in ranking-oriented and variance-explanation metrics across warm start, drug cold start, target cold start, and strict pair cold start settings.

Jun-Wen He · 0 citations
Dataset Open access Jul 2026

VitaGraph: building a knowledge graph for biologically relevant learning tasks

VitaGraph is presented, a comprehensive multi-purpose biological knowledge graph built by integrating and refining multiple public datasets and enabling benchmarking of graph-based models and offering the opportunity to tackle tasks such as drug repurposing, PPI prediction, and side-effect prediction, among others.

Francesco Madeddu, Lucia Testa, Gianluca De Carlo et al. · 0 citations