Skip to content
Preprint

GENIE: Generative Neural Inference for Epidemics

Aug 2026 · 0 citations · 49 references
Mathematics Biology

TL;DR

This work introduces Generative Neural Inference for Epidemics (GENIE), a spatio-temporal ML-based framework for high-resolution forecasting of the burden of respiratory pathogens and demonstrates superior performance across a range of measures.

Abstract

The SARS-CoV-2 pandemic highlighted the ongoing risk infectious diseases pose to society and the value of reliable information on the likely future burden. When forecasting an epidemic at fine spatial resolution, traditionally used mechanistic compartmental model struggle to capture highly complex granular transmission dynamics, resulting in inaccurate and overconfident forecasts. However, detailed Agent-Based Models (ABMs), are challenging to calibrate and are too computationally expensive to use in real-time. Amortized simulation-based inference promises to overcome this difficulty by exploiting the power of machine learning (ML) to perform approximate forecasting at near-real-time using arbitrarily complex models of epidemics. In this work we introduce Generative Neural Inference for Epidemics (GENIE), a spatio-temporal ML-based framework for high-resolution forecasting of the burden of respiratory pathogens. GENIE is designed to reflect two key characteristics of outbreaks: (i) shared biological mechanisms across locations and (ii) location-specific characteristics affecting transmission dynamics. This results in the model architecture having two modules: (i) a Local Infection Encoder - which learns to represent disease dynamics shared across all locations and (ii) a Local Profile Encoder - which learns location-specific representations. Using simulations from a high-resolution spatio-temporal ABM, GENIE is trained to generate samples from an approximate posterior predictive distribution of future epidemic trajectories. Benchmarked against established statistical and ML models, GENIE demonstrates superior performance across a range of measures including the timing and magnitude of peak hospitalisations.

View source

Similar papers

#small language model Preprint Aug 2026

Reservoir: A Large-Scale Simulated Dataset for Training and Evaluating Epidemiological Models

Large-scale, standardized datasets have driven many advances in AI-based scientific modeling, from protein structure prediction to natural language processing. Infectious disease epidemiology is increasingly adopting AI methods for forecasting, surveillance, and outbreak analytics, but the time-series data available to train them remains orders of magnitude smaller than the corpora behind the advances seen in other fields. Because the scope of real-world epidemiological data cannot practically reach the scale needed to train truly large-scale AI methods, simulated data provides a possible alternative. Here we introduce Reservoir, a large open simulator and dataset of realistic epidemic simulations in which every trajectory carries complete ground-truth labels, including quantities that cannot be measured directly in a real outbreak, such as true infection counts, time-varying reproduction numbers, and counterfactual intervention effects. Reservoir is generated by a stochastic simulator with realistic noise and reporting artifacts, together with interventions with configurable timing, compliance, and age-dependent efficacy. The current release contains 500,000 outbreak trajectories spanning one billion simulated days across diverse pathogen characteristics, population structures, and intervention regimes. Reservoir enables counterfactual experiments, surveillance-design studies, and training of epidemic models at a scale real-world datasets cannot provide.

Carson Dudley, Reiden Magdaleno, Marisa Eisenberg · 0 citations
Case report Open access Jul 2026

Retrospective evaluation of variant forecasting models

Variant forecasting is potentially useful for predicting waves of infection and for informing medical countermeasure distribution. To avoid confusion, variant forecasters should distinguish between (1) genomic sequences, which are the data, (2) the taxa (e.g., evolutionary clades or lineages) to which sequences are assigned, and (3) the modeled units formed by aggregating taxa as motivated by real-life epidemiology and practical modeling constraints. Ensembling may be ineffective for variant forecasting because there are few extant classes of variant forecasting models. Rather than rely on a diversity of model types, modelers should consider post hoc model comparison to interrogate and improve the small number of extant models. Most variant forecasting models treat the compositional prevalence of pathogen variants, not the per capita pathogen infection prevalence. This separation is acceptable for the moment but substantially reduces the utility of the resulting forecasts by decoupling variant dynamics from waves of infection. Variant forecasts are typically used to answer decision-oriented questions, like whether a variant will become dominant, while extant scoring metrics focus on a forecast’s ability to make precise estimates day-by-day. Thus, there is a potential disconnect between the performance of models vis-a-vis traditional forecasting scores and their performance in answering the questions that variant forecasts are most often used for.

Thanasi Bakis, Andrew F. Magee, S. Olesen · 0 citations
Open access Jul 2026

Physics-Informed Neural Networks (PINNs) for Parameter Estimation in the SIRS-D Epidemiological Model

Infectious diseases exhibit complex and rapidly evolving transmission dynamics, requiring modeling approaches that can accurately capture these mechanisms. The SIRS-D compartmental model provides a suitable framework, as it incorporates temporary immunity and disease-induced mortality within the epidemic process. Accurate parameter estimation is essential for quantifying the transmission rate, recovery rate, waning immunity rate, and mortality rate, which collectively govern the system behavior. Among existing estimation methods, Physics-Informed Neural Networks (PINNs) offer significant advantages by integrating observational data with the underlying structure of differential equations, thereby preserving physical consistency while maintaining robustness under imperfect data conditions. In this study, PINNs are employed to estimate the parameters of the SIRS-D model using synthetic data generated through the fourth-order Runge–Kutta (RK4) method to ensure stable and consistent numerical solutions. To better represent real-world measurement conditions, 5% noise is added to the synthetic data, introducing realistic variability into the training process. The results demonstrate that PINNs successfully reconstruct the trajectories of S(t), I(t), R(t), and D(t) with low prediction errors. The model achieves MAE values of 0.0065 (S), 0.0067 (I), 0.0208 (R), and 0.0043 (D), with corresponding RMSE values of 0.0090, 0.0074, 0.0253, and 0.0058. Moreover, the estimated parameters closely match the true values, yielding ????????=0.5031, ????=0.0996, ????=0.0095, and ????=0.0149, demonstrating strong parameter identification capability. These findings confirm that PINNs constitute a reliable and accurate framework for analyzing infectious disease dynamics and offer promising potential for extension to more complex epidemiological models and real-world datasets.

Fitri Cahyani, Abdurakhman Abdurakhman, Chyntia Meininda Anjanni · 0 citations
Aug 2026

MechGNN-Epi: Mechanistically Constrained Spatiotemporal Graph Learning for Regional Epidemic Forecasting

This work proposes MechGNN-Epi, a hybrid framework that couples a spatiotemporal graph encoder with a differentiable SIR update that yields epidemiologically constrained trajectories and produces region- and time-indexed parameter proxies that can be inspected as diagnostic signals, while not being guaranteed as causally identifiable mechanistic parameters.

Debashis Chatterjee, Sagnik Acharyya, Subrata Rana · 0 citations
Preprint Jul 2026

Hybrid SINDy-EnKF in Learning Chikungunya Dynamics from Incomplete, Noisy or Partially Observed Data

Current mechanistic models for the transmission dynamics of the Chikungunya virus (CHIKV) rely on uncertain parameters or partially observed data. This limitation challenges the use of theoretical models for understanding and forecasting disease spread. Here we present a hybrid, data-driven model framework that combines Sparse Identification of Nonlinear Dynamics (SINDy) with the Ensemble Kalman Filter (EnKF) for sequential data assimilation. Our numerical experiments show that this approach improves prediction accuracy and provides a good reconstruction of unobserved trajectories under partial observability, a common constraint in real-world epidemiological surveillance. SINDy can be applied to epidemic trajectories, recovering the underlying equations in noise-free conditions. However, standalone SINDy is highly sensitive to noise, leading to spurious terms and poor performance. Hence, we embed the identification procedure within an EnKF framework, which assimilates noisy observations to correct forecast states from the SINDy-derived model and to infer unobserved state variables.

B. A. Afful, Changhong Mou, Luis F. Gordillo · 0 citations
Open access Jul 2026

ADAPTIVE MULTI-LAYER LEARNING FRAMEWORK FOR OUTBREAK PREDICTION

In pandemic situations, it is crucial to analyse outbreaks to identify infections and predict their patterns. Machine Learning models are a powerful tool for outbreak analysis and provide effective insights for pandemic management. So far, most of the work has focused on specific pandemics or outbreaks with similar characteristics. But as history shows, every outbreak has unique traits or symptoms that don't match those in previous studies, and that’s where prediction models often fail. The goal of this study is to provide a framework that is more generic and can be applied to any kind of spread that may turn into an epidemic or pandemic. The major challenge is to find a common framework that can be applied to all pandemics or epidemics. In this work, six pandemic datasets were collected and pre-processed using data cleaning, min-max normalisation, and data augmentation methods, such as Synthetic Minority Over-sampling Technique (SMOTE), to obtain a balanced dataset. Further, a hybrid machine learning model with a three-layer architecture is applied for outbreak prediction. The first layer is an Adaptive Boosting Support Vector Machine (AdaBoost SVM) for sequence modelling. The second layer is K-Nearest Neighbours (KNN) for better categorisation, and the third layer is an optimised Artificial Neural Network (O-ANN) for precise prediction with an accuracy 97.73%, which is better than the standard model.

Artika Singh, Manisha Jailia · 0 citations