Skip to content

BAMPDA: a Matrix Refactoring Framework with Heterogeneous Inference for Potential PTM-disease Association Identification.

Aug 2026 · IEEE transactions on computational biology and bioinformatics · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

BAMPDA is developed, a matrix refactoring framework with heterogeneous inference designed to enhance the inference of PTM-disease associations and rigorous 5-fold cross-validation demonstrates that BAM PDA significantly outperforms five recent baseline methods.

Abstract

Post-translational modifications (PTMs) are modifications of proteins that occur after translation. They exert their influence on human health by altering the properties of proteins. Recent progress in biomedical research has yielded extensive heterogeneous datasets detailing protein PTMs and their disease relevance. These datasets provide a substantial foundation for the development of methods predicting PTM-disease associations. Thus we develop BAMPDA, a matrix refactoring framework with heterogeneous inference designed to enhance the inference of PTM-disease associations. Our approach integrates multi source data encompassing both protein and disease characteristics, leveraging two successively applied matrix based algorithms for association prediction. Furthermore, rigorous 5-fold cross-validation demonstrates that BAM PDA significantly outperforms five recent baseline methods. Finally, we applied BAMPDA to four representative diseases to predict associated proteins and validate the role of PTMs in mediating these disease relationships. The results underscore BAMPDA's robust predictive capability in practical scenarios.

View source

Similar papers

Open access Aug 2026

CLASPP: A unified model for predicting post-translational modifications

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP’s performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Nathan Gravel, Zhongliang Zhou, Ruili Fang et al. · 0 citations
Open access Aug 2026

Biochemically Constrained Multi‐Omics Integration Reveals Protein–Metabolite Dependencies Across Diseases

ABSTRACT Integrating proteomic and metabolomic data is essential for understanding complex diseases, yet current approaches that rely primarily on statistical associations often overlook the structured biochemical relationships between molecular entities and suffer from discriminative instability in small clinical cohorts. Here, we present ProMetNet, a biochemically constrained framework that incorporates pathway‐derived connectivity from the Reactome database into neural network architecture. By encoding protein–metabolite relationships based on reaction topology, ProMetNet models structured cross‐omics dependencies rather than relying solely on statistical correlations, reducing spurious associations while preserving global molecular context and improving robustness in data‐limited settings. Across four heterogeneous disease cohorts, including Alzheimer's disease, type 2 diabetes, COVID‐19, and glioblastoma, ProMetNet consistently outperforms evaluated multi‐omics integration methods, including MOGONET, P‐NET, PEARL, and MOINER, maintaining high discriminative performance under substantial data downsampling. In addition to classification accuracy, the framework prioritizes biologically plausible protein–metabolite dependencies that are not captured by conventional differential or correlation‐based analyses. Importantly, pathway‐level signals identified by ProMetNet demonstrate consistent discriminative performance in independent large‐scale population data from the UK Biobank (N = 47,507), supporting their robustness and generalizability. Together, these results establish ProMetNet as a biologically grounded and interpretable framework for multi‐omics integration, enabling robust identification of structured molecular dependencies across diseases.

Minghui Zhao, Na Zhou, Ruotong Liu et al. · 0 citations
Open access Aug 2026

Interpreting Protein Language Models: high attention sites predict functional regions

The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.

Sophia J. Pribus, Russ B. Altman, Gowri Nayar · 0 citations
Aug 2026

A Novel Graph Transformer Framework for Predicting Drug-Disease Associations with Structural Awareness.

SGTL-DDA is proposed, a novel graph transformer framework designed to incorporate structural information and domain-specific knowledge from heterogeneous biological information networks (HBINs) that successfully identifies both known therapeutics and novel repositioning candidates, supported by molecular docking results and literature evidence.

Bowei Zhao, Hui Zhao, Yu-an Huang et al. · 0 citations
Review Open access Aug 2026

TRACE: A FINE-TUNED BIOMEDICAL LANGUAGE MODEL FOR DIRECTIONALLY INFORMED DRUG REPURPOSING FROM TRANSCRIPTOME-WIDE ASSOCIATION STUDIES

Transcriptome-wide association studies (TWAS) can identify genes where genetically predicted gene expression is associated with disease risk, but translating those signals into therapeutic opportunities remains time-consuming, manual, and difficult to reproduce. We developed TRACE (TWAS-driven Repurposing through AI-assisted Curation of Evidence), a gene- and phenotype-agnostic computational pipeline that accepts a TWAS gene and effect-size direction, normalizes the gene symbol, retrieves FDA-approved drug-gene candidates from four online resources, collects related peer-reviewed literature from PubMed, and uses a fine-tuned biomedical language model to classify whether the literature supports a direct drug-gene relationship, the mechanism of action, and the direction of effect. The pipeline then compares the drug-derived direction with the direction implied by the TWAS effect estimate to rank candidate therapeutic pairs and flag potential drug safety concerns. The local classifier, built on BiomedBERT, was trained using pipeline-derived labels, BioCreative VI ChemProt gold-standard chemical-protein relation examples, and author-reviewed active-learning cases, reaching a held-out macro F1 of 0.809 across three simultaneous classification tasks. We validated the pipeline against a manually curated endometriosis gold standard of 43 drug-gene pairs spanning six TWAS-identified genes, developed through S-PrediXcan analysis of endometriosis GWAS summary statistics, manual querying of four drug-gene interaction databases for each gene, literature review of drug-gene mechanistic evidence, and Mendelian randomization validation of candidate pairs. External validation used two independently published genetically informed drug-repurposing studies in metabolic dysfunction-associated steatotic liver disease (MASLD) and type 2 diabetes (T2D). The pipeline recovered 90.7% of endometriosis pairs, 88.2% of MASLD pairs, and 92.9% of T2D pairs that were present in at least one queried database. Applied to 99 endometriosis-associated TWAS genes, the pipeline identified 1,089 FDA-approved drug-gene pairs, 32 candidate therapeutic pairs, and 77 potential safety concerns, including independent recovery of leuprolide acetate, an established endometriosis therapy. This framework provides a scalable, literature-grounded bridge from TWAS discovery to prioritized therapeutic hypotheses, while preserving uncertainty through manual-review flags and requiring downstream Mendelian randomization, electronic health record-based validation, and experimental follow-up before clinical interpretation.

C. O. Otieno, H. Seagle, A. Akerele et al. · 0 citations
Open access Aug 2026

Large-scale structure prediction of DUF-containing protein-protein interactions

Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.

Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al. · 0 citations