VitaGraph is presented, a comprehensive multi-purpose biological knowledge graph built by integrating and refining multiple public datasets and enabling benchmarking of graph-based models and offering the opportunity to tackle tasks such as drug repurposing, PPI prediction, and side-effect prediction, among others.
Abstract
The complexity of human biology poses ongoing challenges, driving global interdisciplinary research. Artificial intelligence has become a powerful tool in computational biology, where graph data structures model entities like protein–protein interaction (PPI) networks and gene functional networks. These networks support crucial tasks in network medicine, including gene–disease association prediction, drug repurposing, and polypharmacy side-effect analysis. Reliable machine learning predictions require high-quality data. We present VitaGraph, a comprehensive multi-purpose biological knowledge graph built by integrating and refining multiple public datasets. Extending the Drug Repurposing Knowledge Graph, our pipeline: (a) resolves inconsistencies and redundancies, (b) consolidates information from leading public sources, and (c) enriches graph nodes with expressive features such as molecular fingerprints and gene ontologies. Incorporating biologically and chemically meaningful features enhances machine learning models’ ability to learn accurate, structured embedding spaces. The resulting resource offers a coherent, reliable platform to advance computational biology and precision medicine while enabling benchmarking of graph-based models and offering the opportunity to tackle tasks such as drug repurposing, PPI prediction, and side-effect prediction, among others.
SGTL-DDA is proposed, a novel graph transformer framework designed to incorporate structural information and domain-specific knowledge from heterogeneous biological information networks (HBINs) that successfully identifies both known therapeutics and novel repositioning candidates, supported by molecular docking results and literature evidence.
Bowei Zhao, Hui Zhao, Yu-an Huang et al.· IEEE transactions on computa...· 0 citations
Drug repurposing represents a cost-effective strategy to identify novel therapeutic applications for existing pharmaceuticals, circumventing the protracted timelines of traditional drug discovery. While knowledge graph (KG) based methods excel at integrating heterogeneous biomedical data, they often struggle to harmonize high-level domain knowledge with fine-grained molecular mechanisms. We propose KGDDA, a multimodal framework designed for drug-disease association prediction that synergistically integrates KGs with medical ontologies. By leveraging an attention-driven fusion mechanism, KGDDA dynamically merges contextual topological embeddings with ontology-derived priors, enabling the adaptive capture of intricate drug-disease interactions. Extensive evaluations on two benchmark datasets demonstrate that KGDDA consistently outperforms state-of-the-art baselines in both predictive accuracy and generalization. Furthermore, case studies on head and neck cancer and small cell lung cancer validate KGDDA's ability to provide actionable mechanistic insights, highlighting its potential to accelerate therapeutic discovery and precision medicine.
Qichang Zhao, Qiao Ling, Muhammad Habibulla Alamin et al.· IEEE transactions on computa...· 0 citations
External validation on BioKG shows that NEOGRAN remains effective under differences in entity coverage, relation composition, and local topology, supporting its method-level generalizability across knowledge graph sources.
This work proposes an enhanced learning framework that deeply integrates structured logical knowledge within GNN models, and demonstrates that incorporating domain-specific relational knowledge leads to better generalization and robustness compared to standard GNNs.
Kai Hodžić, Gustav Šír· ACM Transactions on Intellig...· 0 citations
This work presents HGRL-PPIS, a novel hierarchical graph representation learning approach for predicting protein-protein interaction sites that achieves superior performance over competing methods on multiple benchmark datasets, enabling more reliable detection of protein-protein binding residues.
This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.