Skip to content

Category

diffusion models

524 papers

#machine learning Open access Apr 2026

Accurate and task-agnostic modeling of enzymatic reactions through multimodal relational learning

Enzymatic reactions play an emerging role in a broad spectrum of scientific and industrial applications. The inherent complexity of enzymes, such as their substrate specificity, conformational flexibility, and the vast diversity of reactions involved, poses substantial challenges for the advanced computational prediction of enzymatic reactions with desirable accuracy. Moreover, existing approaches are mostly tailored for a specific sub-task, such as substrate prediction or binding site annotation, which limits their applicability. In this study, we introduce ERAM, a task-agnostic multimodal learning framework capable of addressing a broad range of downstream applications with both accuracy and efficiency. ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data. In enzyme retrieval tasks, ERAM achieves an improvement of 28.31% in mean average precision compared with the state-of-the-art (SOTA) method, CREEP. In substrate prediction tasks, ERAM outperforms the SOTA method ESP, achieving average improvements of 35.53% and 22.97% in Matthews correlation coefficient across two datasets. Additionally, ERAM exhibits commendable interpretability by assigning higher attention weights to binding sites, resulting in lower false-positive rates (42.36%) and higher overlap scores (70.59%) in the unsupervised binding site prediction task compared to RXNAA Mapper. By learning embeddings of substrates, enzymes, and products within a unified knowledge graph latent space, ERAM demonstrates its potential as a versatile and effective tool for enzyme catalysis research.

Yuansheng Huang, Lanqing Li, Wenjia Qian et al. · 2 citations
#natural language process... Open access Apr 2026

LaMGen: LLM-based 3D molecular generation for multi-target drug design

Multi-target drugs hold great promise for treating complex diseases, yet existing methodologies predominantly rely on ligand-based approaches, which lack sufficient biological context and are often confined to specific target pairs, resulting in limited generalizability. Here, we introduce LaMGen, a general-purpose multi-target drug design framework powered by large language models (LLMs). Built on MTD2025, a dataset comprising over 600,000 quantum-accurate molecular conformations and 700,000 multi-target associations, LaMGen directly yields energy-favorable conformations with quantum-level accuracy. The framework integrates ESM-C protein embeddings, rotation-aware ligand tokens, and a TriCoupleAttention module to capture multi-level target–ligand interactions. Across independent benchmarks, LaMGen outperforms diffusion-based model across multiple properties, generating molecules in an average of 0.44 s, while preserving high conformational plausibility. Retrospective analyses demonstrate that LaMGen not only can reproduce molecules identical to known actives, but also consistently produces structurally novel candidates with conserved core scaffolds and superior binding affinities. Designing effective multi-target therapeutics remains a major challenge, as existing ligand- or protein-centric methods struggle to generate biologically contextualized, spatially valid 3D molecules, particularly for triple-target systems. This study introduces LaMGen, an LLM-powered framework that leverages large-scale protein-ligand data and rotation-aware molecular encoding to rapidly produce chemically plausible multi-target candidates, achieving strong zero-shot generalization, superior molecular quality, and robust performance across dual- and triple-target design tasks.

Qun Su, Qiaolin Gou, Hui Zhang et al. · 1 citation
#machine learning Open access Jul 2026

BBBP-Atlas: Unified Interpretable Modeling of Blood–Brain Barrier Permeability across Small Molecules and Peptides

Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.

Xin Shen, Qun Su, Hao Luo et al. · 0 citations
#machine learning Open access Jun 2026

Targeting the intrinsically disordered AR-NTD through a machine learning-based enhanced sampling workflow

Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.

Kai Zhu, Huating Wang, Jintu Zhang et al. · 0 citations
#diffusion models Open access Aug 2026

Quantum Error Correction Coding and Information Diffusion Simulation

This paper investigates the crucial aspect of quantum error correction (QEC) – the propagation of information during encoding and decoding processes. We propose a novel methodology utilizing Monte Carlo simulation to model the probability distribution of information diffusion within QEC schemes. The core claim is that a systematic understanding of this diffusion, through probabilistic analysis, can significantly improve the design and optimization of QEC protocols, ultimately enhancing the reliability of quantum information transmission. Our approach explicitly simulates entangled qubit states and tracks the probability of information spreading during the encoding and decoding operations. The results provide valuable insights into the limitations and potential improvements of existing QEC strategies. We define key parameters such as the error rate, qubit fidelity, and code distance, and use these to drive our simulation. The simulation outputs are then analyzed to derive effective strategies for mitigating information loss and achieving higher levels of quantum data integrity. This work bridges the gap between theoretical QEC design and practical simulation, paving the way for more robust and efficient quantum communication systems. ---

Jincheng Zhang · 0 citations
#diffusion models Open access Aug 2026

Title: Formal Verification of Generative Models

This paper presents a novel framework for formal verification of generative models, focusing on ensuring their stability and preventing the generation of undesirable outputs. Generative models, such as GANs and diffusion models, are increasingly prevalent in various applications, yet their inherent complexity makes them challenging to verify formally. This research introduces a new system that analyzes probabilistic dependencies within these models, identifies potential failure modes, and constructs a formal verification protocol. The core mechanism involves creating a formal system capable of systematically assessing the model's behavior under various conditions, thereby guaranteeing convergence and mitigating risks associated with unexpected outputs. The paper details the system's architecture, illustrates its application with a specific example, and discusses its potential impact on the field of generative AI.

Jincheng Zhang · 0 citations
#diffusion models Open access Aug 2026

Quantum Error Correction Coding and Information Diffusion Simulation

This paper investigates the crucial aspect of quantum error correction (QEC) – the propagation of information during encoding and decoding processes. We propose a novel methodology utilizing Monte Carlo simulation to model the probability distribution of information diffusion within QEC schemes. The core claim is that a systematic understanding of this diffusion, through probabilistic analysis, can significantly improve the design and optimization of QEC protocols, ultimately enhancing the reliability of quantum information transmission. Our approach explicitly simulates entangled qubit states and tracks the probability of information spreading during the encoding and decoding operations. The results provide valuable insights into the limitations and potential improvements of existing QEC strategies. We define key parameters such as the error rate, qubit fidelity, and code distance, and use these to drive our simulation. The simulation outputs are then analyzed to derive effective strategies for mitigating information loss and achieving higher levels of quantum data integrity. This work bridges the gap between theoretical QEC design and practical simulation, paving the way for more robust and efficient quantum communication systems. ---

Jincheng Zhang · 0 citations

Assessing spatially explicit sensitivities to scenario uncertainty through climate emulation

Comprehensive climate risk assessment requires bridging the gap between the socio-economic detail of integrated assessment models and the physical fidelity of Earth System Models (ESMs). However, the high computational cost of ESMs limit their ability to explore the full range of uncertainty across myriad potential emissions scenarios. Here, we present an impact assessment framework coupling the MIT Emissions Projection and Policy Analysis (EPPA) model with the MIT Earth Sampler, a generative diffusion-based climate emulator. The emulator rapidly generates realizations of spatially- and cross-correlated climate fields at a fraction of the computational cost of an ESM. Benchmarking against a pattern scaling technique utilized by the MIT Integrated Global Systems Model (IGSM) confirms Earth Sampler's ability to reproduce several impact-relevant climate variables. We utilize this generative emulation framework to assess local and regional sensitivities to various emissions scenarios, including an early assessment of the projected ScenarioMIP-CMIP7 protocol. Results show that internal variability masks regional climate outcomes between globally distinct scenarios (e.g., 2°C vs. 1.5°C). Furthermore, we demonstrate that temperature overshoot pathways result in substantially higher cumulative heat stress risks compared to stabilization pathways with similar end-of-century outcomes. Generative climate emulation democratizes access to detailed climate projections, enabling the rapid, probabilistic assessment of compound climate hazards essential for robust adaptation and mitigation planning.

Christopher J. Womack, Shahine Bouabid, Xiang Gao et al. · 0 citations
#diffusion models Open access Aug 2026

Temporal Graph Embedding with Causality-Aware Propagation

Temporal graph embeddings aim to capture the evolving behavior of graphs over time, a crucial task in domains like social network analysis, knowledge graph reasoning, and anomaly detection. However, current graph embedding techniques often treat temporal relationships as simple sequential adjacency updates, neglecting the underlying causal structure that governs how nodes influence each other across time. This paper introduces a novel approach – Causality-Aware Propagation (CAP) – that explicitly models causal relationships within evolving graphs to generate more accurate and informative embeddings. CAP leverages event sequences and domain knowledge to define a causal graph, then employs a modified diffusion process where node representations are propagated based on the learned strength of these causal links. The core idea is to move beyond mere connection propagation to represent the *influence* of connections over time. We demonstrate the effectiveness of CAP through a theoretical analysis and explore its potential applications, establishing a foundational technique for temporal graph representation learning. The method's key contribution lies in its integration of causal inference with graph embedding, offering a more robust and interpretable representation of dynamic graph structures.

Jincheng Zhang · 0 citations
#diffusion models Open access Aug 2026

GSP-3D: generalizable 3D Gaussian Splatting with diffusion policy and world model for long-horizon bimanual manipulation

Introduction Learning robust visuomotor policies for bimanual manipulation remains challenging due to the stringent requirements for precise coordination between arms and the ability to generalize across diverse environmental conditions. Existing diffusion-based policies often suffer from temporally inconsistent action generation, while their reliance on sparse point cloud representations limits structural completeness and fails to capture future scene dynamics, hindering performance in long-horizon tasks. Methods To address these limitations, we introduce GSP-3D, a unified framework that integrates generalizable 3D Gaussian Splatting (3DGS) with diffusion-based policy learning. GSP-3D comprises three key components: (1) a Generalizable Gaussian Regressor (GGR) that predicts 3D Gaussian primitives from a single RGB-D frame in real time; (2) a 3DGS-conditioned diffusion policy that aggregates Gaussians into compact latent representations, replacing point clouds with explicit geometric primitives; and (3) a transformer-based world model that forecasts future Gaussian sets and leverages prediction error as an auxiliary loss to enforce temporal consistency across actions. Results We evaluate GSP-3D on the RoboTwin 2.0 benchmark across a range of bimanual manipulation tasks under both clean and domain-randomized conditions, including variations in lighting, backgrounds, and tabletop distractors. Experimental results show that GSP-3D consistently outperforms existing baseline methods, achieving higher success rates while maintaining minimal computational overhead. Discussion These findings demonstrate that integrating explicit 3D Gaussian representations with diffusion policies offers an efficient and robust solution for long-horizon, temporally coherent bimanual manipulation, effectively addressing the generalization and consistency challenges that limit current approaches.

Xukun Liu, Junhua Huang, Shenggang Wei et al. · 0 citations
#diffusion models Open access Aug 2026

Evaluation of Sugarcane Bagasse as a Biosorbent for Arsenic Removal from Groundwater in Mórrope, Peru

The presence of arsenic in groundwater poses a significant threat to public health, particularly in rural areas lacking access to treatment technologies. In Mórrope, concentrations exceeding regulatory limits have been detected, necessitating sustainable solutions. The objective of this study was to evaluate the efficiency of sugarcane bagasse as a biosorbent material for removing arsenic from contaminated waters in the Mórrope-Lambayeque district. The experimental design employed a 2 × 3 × 3 factorial design, encompassing three variables: adsorbent dosage (1 and 2 g/L), pH (5, 7, and 9), and initial arsenic concentration (2.5, 4.5, and 6.5 mg/L). The bagasse underwent a series of processing steps, including acid washing, drying, and sieving, and was applied in batch tests. The analysis encompassed removal efficiency, adsorption capacity, isotherm adjustment, and kinetics. The system attained a maximum removal rate of 27.43% with a biosorbent dosage of 2 g/L, pH 9, and an initial arsenic concentration of 2.5 mg/L. The Langmuir model (R2 = 0.9926) provided the most suitable description of the adsorption process, indicating the presence of a monolayer. The adsorption kinetics were best described by the pseudo-second-order model (R2 = 0.9727). However, this statistical fit was not interpreted as definitive evidence of an exclusively chemisorption-controlled mechanism, since surface interactions, external mass transfer, and intraparticle diffusion may contribute simultaneously to arsenic uptake. The results indicate that chemically modified sugarcane bagasse exhibits a measurable but limited capacity for arsenic adsorption. However, the maximum removal efficiency of 27.43% is insufficient for direct drinking-water treatment, and the material should not be considered a stand-alone remediation technology under the evaluated conditions. Further surface modification, process optimization, and integration with complementary treatment stages are required before practical implementation can be considered. Beyond its technical feasibility, this approach holds strong social relevance, as it promotes the use of locally available agricultural residues to improve water quality in vulnerable communities, fostering low-cost, community-managed, and environmentally sustainable solutions for safe water access. Although the removal efficiency remains moderate, future research should aim to enhance adsorption performance and evaluate the potential for large-scale implementation in rural water treatment systems.

Alejandro Valencia-Arías, Jorge Eugenio Cabrejos Barriga, Cesar Alberto García Espinoza · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.