Skip to content

Category

diffusion models

479 papers

#artificial intelligence Preprint Aug 2026

DiffSAC: Diffusion-guided Sampling for Consensus-based Robust Estimation

Robust estimation is a core computer vision task frequently tackled using sample consensus. However, traditional methods suffer from inefficient sampling as they struggle to identify effective minimum sets before hypothesis evaluation. To address these challenges, we propose a novel Diffusion-guided Sampling for Consensus-based Robust Estimation (DiffSAC) framework. DiffSAC introduces a diffusion model to learn the distribution of effective minimum sets. It refines the confidence for each data point, indicating whether it belongs to a good minimum set, rather than ranking the data points as in previous work. This significantly reduces the need to process numerous bad sets. To constrain the refinement direction, geometric features are incorporated as conditions within our diffusion model. Consequently, DiffSAC outputs a small number of high-quality minimum sets, enabling identification of the best hypothesis via consensus evaluation. Notably, compared to previous works requiring evaluating over ten thousand hypotheses, DiffSAC achieves state-of-the-art performance with only dozens, significantly boosting efficiency. Extensive experiments across five classic computer vision tasks demonstrate the superiority of DiffSAC. The diffusion model's sampling accelerators enable real-time operation, and DiffSAC can be used as a plug-and-play module to improve existing sample consensus methods.

Chang Nie, Guangming Wang, Zhe Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation

Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.

Lingfeng Yao, Chenpei Huang, Xingke Yang et al. · 0 citations
#diffusion models Open access Aug 2026

Robust unsupervised domain adaptation for medical image segmentation via frequency-conditioned graph diffusion

Cross-domain variability in medical imaging, arising from differences in scanners, acquisition protocols, and patient populations, remains a major challenge for reliable semantic segmentation. Existing unsupervised domain adaptation (UDA) methods predominantly rely on image-level transformations or feature alignment, which often fail to preserve anatomical consistency under large domain shifts. In this work, we propose a novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance. Specifically, we introduce a Frequency-Conditioned Graph Diffusion paradigm, where convolutional features are transformed into anatomical graphs to explicitly capture structural relationships. A latent diffusion process then progressively refines these graph embeddings, guided by frequency-aware contextual cues, enabling robust cross-domain alignment. To further enhance generalization, we integrate structural consistency regularization with adversarial latent alignment, eliminating the need for labeled target data. A dedicated decoder reconstructs dense segmentation maps, while stochastic diffusion sampling provides uncertainty estimates for improved potential clinical reliability. Extensive experiments on multiple public medical imaging benchmarks demonstrate that our method consistently outperforms state-of-the-art UDA approaches, achieving superior segmentation accuracy and robustness under significant domain shifts. These results highlight the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.

Usman Ahmad Usmani, Arunava Roy, Junzo Watada · 0 citations
#diffusion models Review Open access Aug 2026

Identifying Gaps and Future Research Agenda: Key Success Factors of Digital Transformation Adoption

The purpose of this study is to advance knowledge for managers, policymakers, and researchers regarding the important key factors of successful Digital Transformation (DT) adoption, which can shed light on research gaps and help form a research agenda. A systematic literature review was conducted to analyse 26 peer-reviewed journal articles published between 2020 and 2025 from the Scopus, Emerald, Insight, Google Scholar, and ProQuest databases, following the PRISMA guidelines. Four DT adoption models, such as Diffusion of Innovation (DOI), Technology Acceptance Model (TAM), Task-Technology Fit (TTF), and Theory of Planned Behaviour (TPB), have been analysed to understand why people embrace DT either positively or negatively. According to the literature, the success factors for adopting DT have been identified. In addition, it is found that using only one DT adoption model does not ensure success. It is advisable to use an integrated multi-model or multiple frameworks for the theoretical adoption of DT. Using multiple frameworks makes DT adoption easier to understand. The Input-Process Output (IPO) schema allows consolidating the gaps in the present state and setting out the research agenda. The IPO schema can be considered a helpful tool for identifying research gaps and setting the research agenda, as well as for planning, decision-making, and policymaking.

Amando Singun · 0 citations
#diffusion models Open access Aug 2026

Guided protein structure generation for pathway discovery: a showcase for RAF dimerization

Understanding the mechanisms underlying large scale protein conformational changes in signaling pathways is critical for elucidating disease processes and developing targeted therapeutics. However, existing experimental and computational methods struggle to resolve the dynamic ensembles of intermediate states that mediate such transitions, particularly in large biomolecular complexes. Here, we introduce a two-stage generative diffusion modeling framework designed to support pathway discovery in protein complexes, demonstrated using RAF kinase dimerization, a key event for kinase activation and oncogenic signaling. Our approach first generates ultra-coarse-grained structures conditioned on low dimensional descriptors along the monomer-to-dimer transition. It then applies a super-resolution model to recover detailed coarse-grained topologies suitable for molecular simulation. We show that this framework produces physically plausible, diverse, and robust intermediate structures, even for previously unseen interpolated descriptor values. The resulting ensemble enables generation of closely spaced candidate intermediate structures between biophysically distinct states, providing valuable starting points for downstream adaptive sampling and mechanistic studies. Overall, our results highlight the potential of diffusion-based generative models to bridge the gap between static structural data and isolated ensembles, and the dynamic complexity of protein signaling pathways.

Tim Hsu, Konstantia Georgouli, Michael Jones et al. · 0 citations
#diffusion models Review Open access Aug 2026

Public-sector digital transformation in the age of generative AI

Digital transformation (DT) remains central to information systems and public administration scholarship, particularly amid the rapid emergence of generative artificial intelligence (GenAI). This study presents a systematic literature review of 125 peer-reviewed articles published between 2021 and 2026 to synthesise contemporary public-sector DT dynamics. Following the PRISMA 2020 reporting standard and thematic synthesis, the review maps conceptualisations of DT, identifies key organisational, technological and environmental drivers, and examines their implications for public value. The findings indicate that DT is predominantly conceptualised as a sociotechnical and public-value-oriented process shaped primarily by leadership, organisational culture and strategic alignment rather than technological investment alone. Organisational and managerial factors emerge as the most consistent predictors of transformation outcomes across diverse institutional contexts, while technological and environmental conditions influence the pace, direction and unevenness of implementation. Despite the rapid diffusion of AI and GenAI in government practice, only a limited proportion of the literature substantively engages with AI, and fewer studies address GenAI, large language models or foundation models, revealing a widening gap between technological developments and scholarly inquiry. Building on the established organisational-technological-environmental framework, the review identifies four additional explanatory dimensions: citizen co-production, digitally induced administrative burden, street-level administrative reconfiguration, and multidimensional public-value evaluation. The study concludes by identifying priorities for future research on GenAI governance, accountability and equity, while offering practical implications for public managers and policymakers pursuing AI-enabled public-sector transformation.

Gideon Mekonnen Jonathan · 0 citations

Precision Design of Fluorogenic Probes via Orthogonal Tuning of Binding and Photophysics for Isoform-Selective ALDH2 Imaging.

Fluorogenic probes that report enzyme activity are essential for studying biological functions. However, designing them for targets with low catalytic turnover and narrow substrate specificity remains a significant challenge. Here, we present a precision design framework that separates the requirements for sensitivity and selectivity by integrating molecular docking, quantum chemical modeling of fluorogenic mechanisms, and targeted fine-tuning of the probe structures. As a proof of concept, we developed A5, a fluorogenic substrate for aldehyde dehydrogenase 2 (ALDH2) that exhibits high isoform selectivity and a >240-fold signal enhancement over the standard NADH assay. A5 enables quantitative imaging of ALDH2 activity across multiple biological scales─in blood samples, live cells, and intact mouse brains─and supports the identification of small-molecule activators with therapeutic potential in an Alzheimer's disease model. This work establishes a modular strategy for creating activity-based probes tailored to challenging enzymatic targets, with broad applications in precision imaging, drug discovery, and mechanistic biochemistry.

Rongrong Tao, Yu Chen, Taorui Yang et al. · 4 citations

LumiCharge: Spherical Harmonic Convolutional Networks for Atomic Charge Prediction in Drug Discovery.

Atomic charge is crucial in drug design for analyzing reactive sites and interactions between ligands and targets. While quantum mechanical methods offer high accuracy, they are generally computationally costly. Conversely, empirical approaches, while computationally efficient, frequently suffer from lack of precision and generalizability. Recent a number of machine learning-based models have been developed for atomic charge predictions, but they struggle with accurately representing molecular structures and capturing the chemical environments affecting atomic charges, thus limiting their generalization and accuracy. To overcome these limitations, we propose LumiCharge, a novel atomic charge prediction framework that incorporates high-order spherical harmonics convolutions and explicitly models multibody interactions. In constructing this model, we employ a strategy that integrates both high- and low-order information, enhancing its geometric spatial perception capability, which is currently underexplored in the field. Benchmark evaluations demonstrate that LumiCharge outperforms state-of-the-art (SOTA) models by 30%-60% across diverse data sets. Additionally, in cross-scale experiments, LumiCharge demonstrates exceptional extrapolation capability and robustness across molecules of varying sizes, effectively overcoming the limitations imposed by molecular sizes. On an external halogen-containing test set, LumiCharge achieves an RMSE of 0.055e, meeting practical application requirements. Finally, a case study of virtual screening for the androgen receptor (AR) target further validates its outstanding accuracy compared to the OPLS3e force field and other deep learning (DL)-based baseline models, highlighting its exceptional generalization capacity and practical utility in real-world scenarios.

Qun Su, Hui Zhang, Qiaolin Gou et al. · 2 citations

PepBAN: A Deep Learning Framework with Bilinear Attention and Adversarial Learning for Peptide-Protein Interaction Prediction

Accurate prediction of the peptide-protein interaction (PepPI) is crucial for developing peptide-based therapeutics and vaccines. However, this computational task has traditionally faced significant challenges, such as the scarcity of structure data along with the corresponding label of the binding affinity for bound complexes. To address these challenges, we introduce PepBAN, a deep learning framework for modeling PepPI predictions. PepBAN incorporates two technical advancements: (1) adopting the protein language model ESM-2 to characterize proteins and ESM-2 or a graph-based foundation model for peptides without structure data and (2) leveraging the conditional domain adversarial learning to enhance generalization across a broad range of protein targets, especially when there are limited binding data. At the core of PepBAN is a bilinear attention network (BAN) that effectively learns the pattern of pairwise local interactions, enables the identification of key residues participating in the peptide-protein interactions, and offers an intuitive approach to interpret the underlying mechanisms of PepPIs via analyzing attention weights. Our numerical experiments demonstrated that PepBAN outperformed the previous state-of-the-art models across several well-established benchmark studies. Furthermore, we evaluated PepBAN's applicability in predicting cyclic peptide-protein interactions, a task that poses significant challenges due to the presence of noncanonical amino acids. These nonstandard residues require specialized handling, which most existing sequence-based PepPI prediction models did not adequately address, and we adopt an atom-resolved molecular graph approach to process cyclic peptides. Despite this complexity, PepBAN demonstrated a clear advantage by achieving a superior prediction performance and offering a distinct edge in tackling the emerging chemical space of cyclic peptides, which has great potential for novel therapeutic development. In summary, PepBAN serves as a valuable tool for advancing peptide-based drug and therapeutic development.

Shuaiyan Li, Xiaorui Wang, Yuchen Zhu et al. · 2 citations

Discovery of N-(thiazol-2-yl) Furanamide Derivatives as Potent Orally Efficacious AR Antagonists with Low BBB Permeability.

Resistance-conferring mutations in the androgen receptor (AR) ligand-binding pocket (LBP) compromise the effectiveness of clinically approved orthosteric AR antagonists. Targeting the dimerization interface pocket (DIP) of AR presents a promising therapeutic approach. In this study, we report the design and optimization of N-(thiazol-2-yl) furanamide derivatives as novel AR DIP antagonists, among which C13 was the most promising candidate. C13 exhibited excellent AR antagonistic activity (IC50 = 0.010 μM), effectively blocked AR dimerization and nuclear translocation, and demonstrated potent efficacy in several castration-resistant prostate cancer (CRPC) cells. Notably, C13 showed superior efficacy against variant drug-resistant AR mutants, along with favorable metabolic stability, excellent pharmacokinetic properties, and low brain distribution. Furthermore, oral administration of C13 achieved 123.4% tumor growth inhibition in an LNCaP xenograft model without apparent toxicity. As a noncompetitive binder, C13 complements current LBP-targeting AR inhibitors and represents a promising therapy for drug-resistant PCa.

Jinbiao Liao, J. Liao, Yanzhen Yu et al. · 1 citation

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.