Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 12667-12678· 0 citations· 36 references
TL;DR
A text-guided molecular graph generation framework that leverages the structural modeling power of graph diffusion models to achieve both strong alignment with textual descriptions and high-quality molecular structures and a molecule structure consistency loss that explicitly enforces structural coherence during generation, leading to higher-quality and more consistent molecular graphs.
Abstract
Text-guided molecule generation enables controlled molecular design from natural language descriptions and has broad applications in areas such as drug discovery. While recent methods have demonstrated promising capability in generating molecules that align well with textual descriptions, they often overlook the structural properties of the generated graphs. As a result, these approaches struggle to simultaneously ensure consistency with the input text and high structural quality of the generated molecules. In this paper, we propose a text-guided molecular graph generation framework that leverages the structural modeling power of graph diffusion models to achieve both strong alignment with textual descriptions and high-quality molecular structures. However, accomplishing this goal involves several key challenges: 1) how to align graph diffusion models with natural language instructions in order to generate molecular graphs with expected relational semantics from text, 2) how to directly optimize the quality of the generated molecular graphs without sacrificing fine-grained alignment with text-specific details. To tackle these challenges, we introduce Text-guided Conditional Discrete Graph Diffusion (TDGD), a discrete diffusion-based framework for generating molecular graphs from natural language descriptions. Our model incorporates a structure-aware cross-attention mechanism that aligns textual semantics with molecular structures by capturing relational semantics between textual descriptions and molecular structures. In addition, we propose a molecule structure consistency loss that explicitly enforces structural coherence during generation, leading to higher-quality and more consistent molecular graphs. Extensive experiments on ChEBI-20 and L+M-24 datasets demonstrate the effectiveness of our proposed TDGD model.
Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared b...
Dong Xu, Zhang-Fan Yang, Junchuang Cai et al.· IEEE transactions on computa...· 1 citation
Fragment-based drug discovery (FBDD) relies heavily on the design of chemically viable linkers to connect fragments binding to different pocket regions into potent lead molecules. While recent generative models have advanced spatial fragment linking, they frequently produce linkers characterized by high torsional strai...
Kun-Yang Sun, Ying-Ze Wang, Justin Purnomo et al.· Journal of Chemical Informat...· 0 citations
Recently, Large Language Models (LLMs) have become the dominant paradigm for molecular editing due to their strong generalization capabilities across diverse tasks. However, treating molecules as 1D text strings (SMILES) introduces significant challenges in controllability and structural validity. In this paper, we que...
Jiajun Yu, Zhihao Wu, Yizhen Zheng et al.· Proceedings of the 32nd ACM...· 0 citations
This is the first method to expose GNN-derived attributions to an LLM as evidence for property prediction, and achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task.
Junwoo Park, Minyoung Shin, C. Lee et al.· 0 citations
This work introduces \textbf{MolEmb}, a lightweight framework that adapts MLLMs by aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, and finds that context-aware molecular embedding is primarily a data property of the supervision.
Xinjian Zhao, Xiang-Ru Jian, Yao-Yao Xu et al.· 2 citations
Results indicate that FragSyn, through the synergistic design of fragment-level representation and cellular context awareness, provides an effective and interpretable new approach to synergistic drug combination prediction.
Li-Feng Shao, Jianqiang Sun, Hong-Zhan Ma et al.· Journal of Chemical Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.