Skip to content

DDI-MMAF: Multi-modal affine fusion of visual and semantic representations for anticancer drug synergy prediction.

Jul 2026 · European journal of medicinal chemistry · Vol 317, pp. 119134 · 0 citations · 45 references
Medicine

TL;DR

DDI-MMAF is proposed, a lightweight cross-modal framework that avoids using explicit high-dimensional omics profiles as direct model inputs and enables accurate synergy prediction from raw minimalist inputs, offering a practical and efficient computational solution for cost-effective drug combination discovery.

Abstract

Predicting anticancer drug synergy is pivotal for personalizing combination therapies; however, existing deep learning models often rely heavily on complex, high-dimensional multi-omics data and precomputed molecular properties. Such dependence increases data acquisition barriers and limits model applicability in resource-constrained or rapid screening scenarios. In this study, we propose DDI-MMAF, a lightweight cross-modal framework that avoids using explicit high-dimensional omics profiles as direct model inputs. It utilizes a minimalist input protocol consisting of drug SMILES sequences and verbalized biological context, encompassing cell line names and their corresponding tissue origins. The architecture integrates a domain-specific semantic encoder to extract coarse-grained biomedical semantic priors from cell-line nomenclature and a deep residual visual network to capture hierarchical spatial features from molecular images. The core innovation lies in a multi-modal affine fusion mechanism that dynamically modulates molecular visual features conditioned on biological semantic embeddings. Systematic evaluations demonstrate that despite its simplified inputs, the model achieves a ROC AUC of 0.934 on benchmark datasets, showing the best performance against methods that utilize explicit omics information or handcrafted molecular descriptors. Furthermore, our approach maintains robust performance under the internal scaffold-split setting, while also achieving competitive performance compared with the evaluated baselines on the independent AstraZeneca blind test set. Overall, this research demonstrates that effective semantic-guided modulation enables accurate synergy prediction from raw minimalist inputs, offering a practical and efficient computational solution for cost-effective drug combination discovery.

View source

Similar papers

Open access Aug 2026

Multimodal contrastive learning for integrating molecular representations and cellular phenotypes in drug-target interaction prediction

Abstract Motivation Accurate prediction of drug-target interactions (DTIs) is fundamental to drug discovery and mechanistic understanding. While deep learning has advanced computational DTI prediction, most existing methods rely primarily on molecular structural representations, including drug structures and protein sequences, while overlooking cellular phenotypes that reflect downstream biological effects. Cell Painting enables high-content morphological profiling that captures systems-level responses to chemical and genetic perturbations but remains underutilized in DTI modeling. Integrating molecular information with cellular phenotypes offers an opportunity to improve both predictive performance and biological interpretability. Results We propose a two-stage contrastive learning framework integrating drug structures, protein sequences, and Cell Painting morphological profiles into a unified embedding space. Stage 1 learns modality-specific representations independently from structure-based and image-based data; Stage 2 aligns these via multi-positive contrastive learning to bridge molecular structural information with cellular phenotypes. Cross-modal retrieval achieves median Recall@10 values of 0.77 (random split) and 0.33 (scaffold split), outperforming bilinear and random baselines. In external DTI prediction on the BIOSNAP dataset, our model achieves an AUC of 0.92 with image-based representations and 0.90 under structure-only settings, surpassing existing methods. Model interpretation via integrated gradients reveals pathway-specific morphological signatures associated with drug targets, providing biologically interpretable insights into drug mechanisms. Availability https://github.com/YJRubyLai/Unified-DTI

Ying-Ju Lai, Tianyuzhou Liang, Po-Yuan Chen et al. · 0 citations
Jul 2026

TextDTI: A Multimodal Context Representation Learning Framework for Drug-Target Interaction Prediction.

This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.

Jiaqi Deng, Senyu Tang, Jijun Tang et al. · 0 citations
Open access Jul 2026

A multi-view feature fusion framework with interpretable graph convolution for predicting microbe-drug associations.

IDEAL (Interpretability-Driven Evolvable Attentive Learning for Microbe-Drug Association) is proposed, a multi-view framework that integrates drug network topological attributes, BERT-encoded drug semantics, drug fingerprints, microbe genome sequence attributes, BERT-encoded microbe semantics, and microbe metabolic pathway attributes.

Lisha Zhou · 0 citations
Aug 2026

CMAF-DDI: A Knowledge-Enhanced Cross-Modal Fusion Method Leveraging Protein Representation for Multi-Class Drug-Drug Interactions.

Accurate prediction of drug-drug interactions (DDIs) is crucial for medication safety and personalized treatment. Most existing methods primarily exploit molecular graphs or biomedical knowledge graphs, while target protein sequence information is often underused. This paper proposes CMAF-DDI, a multi-class DDI prediction framework that integrates protein sequence features, molecular graph features, and knowledge graph features. CMAF-DDI contains a bi-level cross-modal fusion module: an Attention Fusion (AF) level that models global dependencies among modalities using multi-head attention, and a Triple-feature Product Fusion (TPF) level that captures high-order cross-modal co-activation after projecting all modalities into a shared latent space. Experimental results on DrugBank and DRKG show that CMAF-DDI improves multi-class DDI prediction compared with representative graph-based and multi-source fusion baselines. We further provide ablation, hyperparameter sensitivity, controlled protein perturbation, representation, and case-level analyses to examine the contribution and behavior of protein-enhanced cross-modal fusion.

Hengpeng Zhao, Xiaoli Lin, Jun Pang et al. · 0 citations