Skip to content

Reservoir: A Large-Scale Simulated Dataset for Training and Evaluating Epidemiological Models

Aug 2026 · 0 citations · 42 references
Biology

Abstract

Large-scale, standardized datasets have driven many advances in AI-based scientific modeling, from protein structure prediction to natural language processing. Infectious disease epidemiology is increasingly adopting AI methods for forecasting, surveillance, and outbreak analytics, but the time-series data available to train them remains orders of magnitude smaller than the corpora behind the advances seen in other fields. Because the scope of real-world epidemiological data cannot practically reach the scale needed to train truly large-scale AI methods, simulated data provides a possible alternative. Here we introduce Reservoir, a large open simulator and dataset of realistic epidemic simulations in which every trajectory carries complete ground-truth labels, including quantities that cannot be measured directly in a real outbreak, such as true infection counts, time-varying reproduction numbers, and counterfactual intervention effects. Reservoir is generated by a stochastic simulator with realistic noise and reporting artifacts, together with interventions with configurable timing, compliance, and age-dependent efficacy. The current release contains 500,000 outbreak trajectories spanning one billion simulated days across diverse pathogen characteristics, population structures, and intervention regimes. Reservoir enables counterfactual experiments, surveillance-design studies, and training of epidemic models at a scale real-world datasets cannot provide.

View source

Similar papers

Preprint Aug 2026

GENIE: Generative Neural Inference for Epidemics

This work introduces Generative Neural Inference for Epidemics (GENIE), a spatio-temporal ML-based framework for high-resolution forecasting of the burden of respiratory pathogens and demonstrates superior performance across a range of measures.

Laura M Guzman-Rincon, George R.E. Bradley, Joel Kandiah et al. · 0 citations
Review Open access Aug 2026

Epidemic Modelling in Infectious Disease Dynamics: A Critical Integrative Review of Statistical, Mathematical and Machine-Learning Approaches

Epidemic modelling increasingly combines routine surveillance, mechanistic transmission theory and data-intensive learning, yet the resulting approaches are often compared as though they estimate the same quantities and serve the same decisions. This critical narrative review evaluates statistical, mathematical, machine-learning and hybrid approaches to infectious disease dynamics, with emphasis on inferential purpose, data-generating and observation processes, uncertainty, validation, interpretability and public-health use. Peer-reviewed literature published from January 2000 to 31 May 2026 was identified through biomedical, multidisciplinary and computing-oriented scholarly sources, supplemented by citation searching and verification against authoritative bibliographic records. Foundational earlier papers were retained when necessary. The synthesis indicates that no model class is uniformly superior. Statistical surveillance and time-series models are often efficient for anomaly detection, nowcasting and short-horizon forecasting, but their parameters rarely support intervention counterfactuals without additional causal structure. Mechanistic compartmental, network, spatial and agent-based models make transmission assumptions explicit and can represent intervention pathways, although structural misspecification, weak identifiability and mismatch between latent infections and observed reports can dominate their uncertainty. Machine-learning models can extract nonlinear and high-dimensional patterns from heterogeneous data, but apparent accuracy is vulnerable to temporal or spatial leakage, changing surveillance systems, distribution shift, weak probabilistic calibration and limited causal meaning. Hybrid models can combine epidemiological constraints with flexible learning and data assimilation; their interpretability nevertheless depends on identifiable parameters, biologically coherent architecture and validation beyond the setting used for training. Across paradigms, the observation process, target definition and decision horizon are as consequential as model form. Credible use therefore requires question-first model selection, explicit separation of forecasts from scenarios, rolling and geographically external validation, calibrated uncertainty, versioned data and code, and transparent communication of assumptions. Progress will depend less on greater complexity alone than on prospective benchmarking, identifiable hybridisation, behaviour-aware causal designs, multimodal surveillance, equitable data systems and operational model governance.

Hamid H. Hussien, Muhammed Aljifri, Nuha Hassan Hagabdulla et al. · 0 citations
Case report Open access Jul 2026

Retrospective evaluation of variant forecasting models

Variant forecasting is potentially useful for predicting waves of infection and for informing medical countermeasure distribution. To avoid confusion, variant forecasters should distinguish between (1) genomic sequences, which are the data, (2) the taxa (e.g., evolutionary clades or lineages) to which sequences are assigned, and (3) the modeled units formed by aggregating taxa as motivated by real-life epidemiology and practical modeling constraints. Ensembling may be ineffective for variant forecasting because there are few extant classes of variant forecasting models. Rather than rely on a diversity of model types, modelers should consider post hoc model comparison to interrogate and improve the small number of extant models. Most variant forecasting models treat the compositional prevalence of pathogen variants, not the per capita pathogen infection prevalence. This separation is acceptable for the moment but substantially reduces the utility of the resulting forecasts by decoupling variant dynamics from waves of infection. Variant forecasts are typically used to answer decision-oriented questions, like whether a variant will become dominant, while extant scoring metrics focus on a forecast’s ability to make precise estimates day-by-day. Thus, there is a potential disconnect between the performance of models vis-a-vis traditional forecasting scores and their performance in answering the questions that variant forecasts are most often used for.

Thanasi Bakis, Andrew F. Magee, S. Olesen · 0 citations
Open access Aug 2026

Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data

The validation of methods is an integral part of statistical research, defining conditions under which methods yield reliable results. Empirical validation requires a solid data basis to control and manage relevant characteristics like sample size, dimensionality, and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where privacy regulations restrict availability. For this reason, synthetic data is an effective alternative for method validation. However, generating synthetic data is demanding when it must precisely mirror complex dependence structures while simultaneously controlling specific target characteristics. We address the medical context of co-occurring diseases, where symptoms may overlap or conflict. We propose a four-step framework to generate synthetic data for the simulation-based validation of statistical methods. The framework involves: (I) generating patient covariates; (II) connecting this information to predictors for single or joint disease occurrence; (III) transforming predictors into disease probabilities or scores; and (IV) converting these into disease occurrences. Each step offers several alternatives for modeling the overall dependence structure. We apply our framework to a case study of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. By employing five combinations of methodological alternatives, we evaluate the approaches’ ability to achieve target characteristics and demonstrate their specific strengths and weaknesses. Matching the data-generating process with the estimation method allows for the successful recovery of input information, such as coefficients and correlations. Target properties like disease prevalence and associations are achieved to varying degrees depending on the methods used. While the proposed theory-driven framework is broadly applicable beyond the specific medical use case, it relies on careful, domain-informed parameter curation to generate meaningful synthetic datasets. Its flexible, adjustable input settings enable researchers to tailor data generation to their precise methodological requirements, providing a controlled basis for simulation-based validation without implying direct clinical inference.

H. Marchi, S. Schmiegel, T. Schamberger et al. · 1 citation
Preprint Jul 2026

Fast, Frequentist Estimation of Epidemic Reproduction Numbers

The effective reproduction number $R_t$ is one of the most important indicators of epidemic dynamics. Estimating $R_t$, typically from case reports or hospitalization counts, poses a challenging inverse problem. One key issue is lag: $R_t$ acts at the moment of transmission, while the data it generates surface days later. To handle this delay and infer recent infections in real time, popular methods take a Bayesian approach, which can be slow and sensitive to prior specification. As an alternative, we propose ConvRt, a frequentist method for retrospective and real-time estimation. ConvRt deconvolves latent infections and then estimates $R_t$ with successive penalized-likelihood steps, using a spline basis to model smooth curves. Across both stylized and data-driven simulations, we demonstrate favorable performance in point estimation, uncertainty quantification, and runtime. Moreover, by untangling smoothness from future projections, ConvRt enables researchers to assess which qualitative narratives about $R_t$ the data support.

Jeremy Goldwasser, R. Tibshirani, A. Bilinski · 0 citations
Open access Jul 2026

Fine-tuned large language models enhance influenza forecasting.

Influenza-like illness (ILI) remains a persistent global health challenge, necessitating accurate forecasting tools for timely public health response. This study systematically benchmarks fine-tuned large language models (LLMs), e.g., Llama2 and GPT2, for influenza surveillance forecasting in data-limited time-series settings. We develop a lightweight fine-tuning framework that adapts pre-trained LLMs using compact embedding and prediction layers and evaluate it on seven weekly aggregated real-world surveillance datasets. Despite sample sizes of only ∼523 time points per region and the absence of cloud-based data transfer, fine-tuned LLMs consistently outperform SARIMA, LSTM, PatchTST, CoVTransformer, FEDformer, Time-LLM, and GPT4TS in both accuracy and stability, especially for long-term forecasts across diverse geographic settings. Even in zero-shot settings, pre-trained LLMs capture broad epidemic trends with performance comparable to SARIMA. These findings establish fine-tuned LLMs as efficient and robust forecasting tools suitable for privacy-sensitive, data-scarce public health applications.

Chenxi Li, Wenjing Gao, Qiqiao Zhang et al. · 0 citations

Related blog posts