Skip to content

Category

diffusion models

524 papers

#diffusion models Open access Aug 2026

Weak-Form Learning for Mean-Field Partial Differential Equations: An Application to Insect Movement

Abstract. Insect populations subject to epidemic infection, predation, and anisotropic environmental conditions may exhibit preferential movement patterns. Given the virtually random nature of the exogenous factors driving these patterns over short timescales, individual insect trajectories often resemble overdamped stochastic processes. Consequently, data-driven modeling approaches designed to learn effective mean-field governing equations from observed insect populations serve as ideal tools for understanding and predicting such behavior. In this work, we extend existing weak-form equation learning techniques to learn effective partial differential equation models for lepidopteran larval population movement directly from sparse experimental data. We demonstrate the utility of the method on an experimental dataset obtained in simulated agricultural conditions. Relevance to Life Sciences. Understanding dispersal dynamics of crop and silvicultural pests can lead to improved forecasting of outbreak intensity and location, which results in better pest management. As a case study, we demonstrate the utility of our modeling methodology on a sparse experimental dataset consisting of position measurements of fall armyworms ( Spodoptera frugiperda). The data were obtained in simulated agricultural conditions with varied resource quality and larval infection status with respect to a natural pathogen. Using our weak-form modeling methodology, we characterize the dominant mechanisms in the dispersal dynamics and provide quantitative estimates of the effective diffusion rates. Our results indicate that the dispersal dynamics are primary diffusive, although nonnegligible contributions arise from a nonuniform plant resource distribution. Mathematical Content. Galerkin equation learning methods, such as the Weak-form Sparse Identification of Nonlinear Dynamics (WSINDy) algorithm, have recently proven useful for identifying mean-field governing equations from interacting particle data within several biological contexts. In this work, we adapt the WSINDy algorithm, coupled with kernel density estimation, to learn effective partial differential equation models for lepidopteran larval population movement directly from sparse experimental data. In particular, our method characterizes effective external and interaction potentials, as well as diffusive terms, in nonlinear Fokker–Planck models arising from a system of McKean–Vlasov stochastic ordinary differential equations.

Seth Minor, Bret D. Elderd, Benjamin L. Allen et al. · 0 citations
#diffusion models Open access Aug 2026

Generative Distribution Prediction: A Unified Approach to Multimodal Learning

Abstract Accurate prediction for multimodal data—including tabular, textual, and visual inputs or outputs—is essential for advancing analytics across diverse application domains. Existing methods often struggle to integrate heterogeneous data types while maintaining strong predictive performance. We introduce Generative Distribution Prediction (GDP), a model-agnostic framework that leverages high-fidelity multimodal synthetic data generated from the conditional distribution of interest, such as via conditional diffusion models, to enhance prediction across both structured and unstructured modalities. GDP is compatible with any expressive generative model and naturally supports transfer learning for domain adaptation. We provide a rigorous theoretical foundation for GDP, establishing statistical guarantees on its predictive accuracy when diffusion models serve as the generative backbone. By estimating the underlying data-generating distribution and enabling loss-adapted risk minimization, GDP delivers accurate point predictions in broad multimodal settings. We empirically validate GDP on a range of supervised learning tasks, including adaptive quantile regression, modal regression, tabular prediction, image captioning, and question answering, demonstrating its versatility and effectiveness across domains.

Xinyu Tian, Xiaotong Shen · 0 citations
#diffusion models Open access Sep 2026

Improved Wasserstein Estimates for Jump-Diffusion Perturbations under Semigroup Smoothing

We study terminal perturbation bounds for jump-diffusion processes in the Wasserstein distance $W_1$. In the finite-variance setting, let $\nu_0$ and $\nu_1$ denote the jump kernels and define$$\theta_\nu:=\int_0^Td_{\mathrm{FM}}\left(y^2\nu_0(t,\mathrm{d}y),y^2\nu_1(t,\mathrm{d}y)\right)\mathrm{d}t.$$Here $d_{\mathrm{FM}}$ denotes the Fortet--Mourier distance. We prove that the rate $W_1=O(\theta_\nu^{1/3})$ is optimal in the degenerate class. If $r_\nu(t)$ denotes the instantaneous jump discrepancy and $r_\nu\in L^q(0,T)$, with $1\leq q<\infty$, one-sided smoothing improves the exponent to $q/(q+2)$. For additive processes, a two-sided Duhamel argument removes the terminal singularity and yields the optimal linear rate $W_1=O(\theta_\nu)$ under uniform Gaussian smoothing. We also extend the method to multidimensional additive processes, finite-rank state-dependent kernel perturbations, and infinite-variance models. In the latter case, for stable smoothing of order $\beta\in(1,2)$, the one-sided exponent is $1/(3-\beta+\beta/q)$, while two-sided smoothing again gives a linear estimate.

Alexandre Autran · 0 citations
#diffusion models Open access Sep 2026

Improved Wasserstein Estimates for Jump-Diffusion Perturbations under Semigroup Smoothing

We study terminal perturbation bounds for jump-diffusion processes in the Wasserstein distance $W_1$. In the finite-variance setting, let $\nu_0$ and $\nu_1$ denote the jump kernels and define$$\theta_\nu:=\int_0^Td_{\mathrm{FM}}\left(y^2\nu_0(t,\mathrm{d}y),y^2\nu_1(t,\mathrm{d}y)\right)\mathrm{d}t.$$Here $d_{\mathrm{FM}}$ denotes the Fortet--Mourier distance. We prove that the rate $W_1=O(\theta_\nu^{1/3})$ is optimal in the degenerate class. If $r_\nu(t)$ denotes the instantaneous jump discrepancy and $r_\nu\in L^q(0,T)$, with $1\leq q<\infty$, one-sided smoothing improves the exponent to $q/(q+2)$. For additive processes, a two-sided Duhamel argument removes the terminal singularity and yields the optimal linear rate $W_1=O(\theta_\nu)$ under uniform Gaussian smoothing. We also extend the method to multidimensional additive processes, finite-rank state-dependent kernel perturbations, and infinite-variance models. In the latter case, for stable smoothing of order $\beta\in(1,2)$, the one-sided exponent is $1/(3-\beta+\beta/q)$, while two-sided smoothing again gives a linear estimate.

Alexandre Autran · 0 citations
#diffusion models Open access Sep 2026

Paper UAP-K: Temporal and Spectral Structure in an Instrumented Ball-Lightning Record Public-Data Reconstruction, Mechanism Identifiability, and a Prospective Holosphere Corridor Tes

Ball lightning is a long-reported atmospheric phenomenon for which rare instrumented records remain especially valuable. This paper examines one reported natural event that followed a cloud-to-ground lightning strike and was recorded by two slitless spectrographs at an approximate range of 0.9 kilometers. Publications issued in 2014, 2018, 2022, and 2025 are successive analyses of this single event rather than independent replications. The present study reconstructs the public numerical products deposited for the 2025 analysis and determines whether those products can adjudicate a history-dependent Holosphere transport-corridor mechanism. The reviewed program verifies all ten deposited source files against a frozen byte-count and SHA-256 ledger, checks the expected embedded legend tokens and deposited column-order consistency used for the principal species assignments, reparses five Origin project files into 67 numerical datasets, and inventories 64 cropped spectral strips. The restricted parser does not independently decode the complete Origin graph-object binding. All twelve fixed validation checks pass. The reconstructed mean radiation-power densities agree within one percent with the source-paper approximations: about 374 million watts per square meter for neutral oxygen, 30.3 million watts per square meter for neutral silicon, and 25.7 million watts per square meter for neutral iron. The line-intensity table has unequal data availability, with 31 observed neutral-oxygen values, 22 neutral-nitrogen values, and 32 values each for neutral silicon and neutral iron. Both the available-observation comparison and an identical-row comparison restricted to the 22 rows containing all four lines show the same qualitative pattern: the oxygen and nitrogen emissions are much more variable than the silicon and iron emissions. A robust median-based sensitivity analysis preserves this ordering. The supported result is therefore a descriptive contrast between the dispersion of the air-related and soil-related spectral lines, not the identification of two unique physical states or mechanisms. A post-inspection gap rule divides the deposited neutral-oxygen power values into nine contiguous data runs with a median run-start spacing of 10.3335 milliseconds, corresponding numerically to 96.77 cycles per second. The nine-run result persists for gap thresholds from one to five milliseconds, while a half-millisecond threshold produces eleven runs. These are runs in a processed data deposit, not independently detected physical pulse onsets in a complete raw waveform. No physical period, natural constant, or Holosphere frequency is inferred. Paper K6 defines corridor memory as a bounded spatial field produced specifically by successful admissible transport and modified by leakage and diffusion. The public ball-lightning packet does not measure the three-dimensional parent-lightning channel, a synchronized trajectory of the emitting region, local electric and magnetic fields, channel temperature, electron density, conductivity, space charge, or an independent event-level holdout. A generic source-and-decay memory recurrence is not specifically Holosphere evidence when an ordinary electrical-discharge model can use the same recurrence and observational freedom. A distinct corridor test requires measured source and support variables and predictive improvement beyond ordinary lightning-channel and plasma memory. Paper UAP-K therefore supports reproducible public-data reconstruction and a robust descriptive spectral-dispersion contrast. Spatial corridor reuse remains not adjudicated, and this event does not test or validate a Holosphere mechanism. The paper instead supplies a prospective model ladder, control structure, instrumentation requirements, and theory-modification rules for future multi-event lightning and ball-lightning measurements.

Michael Sarnowski · 0 citations
#diffusion models Open access Sep 2026

Paper UAP-K: Temporal and Spectral Structure in an Instrumented Ball-Lightning Record Public-Data Reconstruction, Mechanism Identifiability, and a Prospective Holosphere Corridor Tes

Ball lightning is a long-reported atmospheric phenomenon for which rare instrumented records remain especially valuable. This paper examines one reported natural event that followed a cloud-to-ground lightning strike and was recorded by two slitless spectrographs at an approximate range of 0.9 kilometers. Publications issued in 2014, 2018, 2022, and 2025 are successive analyses of this single event rather than independent replications. The present study reconstructs the public numerical products deposited for the 2025 analysis and determines whether those products can adjudicate a history-dependent Holosphere transport-corridor mechanism. The reviewed program verifies all ten deposited source files against a frozen byte-count and SHA-256 ledger, checks the expected embedded legend tokens and deposited column-order consistency used for the principal species assignments, reparses five Origin project files into 67 numerical datasets, and inventories 64 cropped spectral strips. The restricted parser does not independently decode the complete Origin graph-object binding. All twelve fixed validation checks pass. The reconstructed mean radiation-power densities agree within one percent with the source-paper approximations: about 374 million watts per square meter for neutral oxygen, 30.3 million watts per square meter for neutral silicon, and 25.7 million watts per square meter for neutral iron. The line-intensity table has unequal data availability, with 31 observed neutral-oxygen values, 22 neutral-nitrogen values, and 32 values each for neutral silicon and neutral iron. Both the available-observation comparison and an identical-row comparison restricted to the 22 rows containing all four lines show the same qualitative pattern: the oxygen and nitrogen emissions are much more variable than the silicon and iron emissions. A robust median-based sensitivity analysis preserves this ordering. The supported result is therefore a descriptive contrast between the dispersion of the air-related and soil-related spectral lines, not the identification of two unique physical states or mechanisms. A post-inspection gap rule divides the deposited neutral-oxygen power values into nine contiguous data runs with a median run-start spacing of 10.3335 milliseconds, corresponding numerically to 96.77 cycles per second. The nine-run result persists for gap thresholds from one to five milliseconds, while a half-millisecond threshold produces eleven runs. These are runs in a processed data deposit, not independently detected physical pulse onsets in a complete raw waveform. No physical period, natural constant, or Holosphere frequency is inferred. Paper K6 defines corridor memory as a bounded spatial field produced specifically by successful admissible transport and modified by leakage and diffusion. The public ball-lightning packet does not measure the three-dimensional parent-lightning channel, a synchronized trajectory of the emitting region, local electric and magnetic fields, channel temperature, electron density, conductivity, space charge, or an independent event-level holdout. A generic source-and-decay memory recurrence is not specifically Holosphere evidence when an ordinary electrical-discharge model can use the same recurrence and observational freedom. A distinct corridor test requires measured source and support variables and predictive improvement beyond ordinary lightning-channel and plasma memory. Paper UAP-K therefore supports reproducible public-data reconstruction and a robust descriptive spectral-dispersion contrast. Spatial corridor reuse remains not adjudicated, and this event does not test or validate a Holosphere mechanism. The paper instead supplies a prospective model ladder, control structure, instrumentation requirements, and theory-modification rules for future multi-event lightning and ball-lightning measurements.

Michael Sarnowski · 0 citations
#diffusion models Open access Sep 2026

Deep Learning-Driven Generative Layout Design Model for Sensory Garden Environment Design

The spatial layout design of sensory gardens involves the collaborative optimization of multiple design decisions. Traditional methods rely on manual experience and are inefficient in scheme exploration. This study proposes a diffusion generation model that integrates design grammar rules and conditional constraints to achieve end‑to‑end generation from land‑use conditions to complete layout schemes. A multi‑scale spatial feature encoder is designed to capture the spatial skeleton and element combination patterns of the garden through global‑ and local‑scale hierarchical coding and cross‑scale attention fusion. A conditional embedding module is constructed, in which hard constraints such as the land‑use red line and setback distance are encoded as differentiable vectors, and soft constraints such as visual permeability and spatial enclosure are encoded as differentiable vectors. A graph convolutional network is used to explicitly model the adjacency, inclusion, and axis relationships among design elements, and the topological prior is injected into the denoising process. The U‑Net backbone is improved, and a dual‑channel attention mechanism and a progressive refinement strategy are introduced to enhance structural awareness and boundary accuracy. Experiments are carried out on a dataset containing 200 cases, and ten indicators such as hard constraint violation rate, Fréchet distance, and spatial enclosure are used for evaluation. The results show that the hard constraint violation rate of this model is 2.3%, which is 84% lower than that of the standard diffusion model. The Fréchet distance is 8.7, and the spatial enclosure is 59.3%. All ten indicators outperform the four baseline methods such as LayoutGAN and LayoutVAE. The ablation experiment confirms the independent contribution of each core component, and the cross‑shape generalization and small‑sample experiments verify the adaptability and data efficiency of the model, which provided an effective method support for generative landscape design.

Jing Zhang · 0 citations
#generative ai Open access Aug 2026

From Redaction to Restoration: Deep Learning for Medical Image Deidentification and Reconstruction.

Removing patient-identifying information from medical images is a prerequisite for sharing image data directly, as in public dataset release and open benchmarks, where the images themselves, rather than only model updates must leave the originating institution. However, many methods currently used for de-identification, e.g., cropping or blacking out image regions to eliminate burned-in text, can have negative effects on downstream image analysis tasks because of removal of relevant but non-identifiable information. This work presents an end-to-end deep learning framework for transforming raw clinical image volumes into de-identified, analysis-ready datasets without compromising downstream utility. The methodology developed and tested in this work first detects and redacts regions likely to contain protected health information (PHI), such as burned-in text and metadata, and then uses a generative deep learning model to inpaint the redacted areas with anatomically and imaging-plausible content. The proposed pipeline leverages a lightweight hybrid architecture, combining CRNN-based redaction with a latent-diffusion inpainting restoration module (Stable Diffusion 2). We evaluate the approach using both privacy-oriented metrics, which quantify residual PHI and success of redaction, and image-quality and task-based metrics, which assess the fidelity of restored volumes for representative deep learning applications. The binary mask performance shows strong overall PHI identification (F1 score = 0.891 ± 0.037, recall of 0.912 ± 0.053, and precision of 0.875 ± 0.058, indicating accurate localization of PHI-containing regions with few missed detections or false-positive redactions. Downstream anatomy identification (segmentation) tasks remain markedly similar across after applying varied inpainting strategies (Diffusion with/without context and Telea with Dice ranging from 0.936 to 0.948 on one dataset and 0.955 to 0.959 on another. Our results suggest that the proposed method yields de-identified medical images that are visually coherent, maintaining fidelity for downstream models and clinical tasks, while substantially reducing the risk of patient re-identification. By automating de-identification and image reconstruction within a single workflow and disseminating large-scale medical imaging collections, thereby lowering a key barrier to data sharing and multi-institutional collaboration in medical imaging AI.

Adrienne Kline, A. Gaonkar, D. Pittman et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.