Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data
Abstract
The validation of methods is an integral part of statistical research, defining conditions under which methods yield reliable results. Empirical validation requires a solid data basis to control and manage relevant characteristics like sample size, dimensionality, and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where privacy regulations restrict availability. For this reason, synthetic data is an effective alternative for method validation. However, generating synthetic data is demanding when it must precisely mirror complex dependence structures while simultaneously controlling specific target characteristics. We address the medical context of co-occurring diseases, where symptoms may overlap or conflict. We propose a four-step framework to generate synthetic data for the simulation-based validation of statistical methods. The framework involves: (I) generating patient covariates; (II) connecting this information to predictors for single or joint disease occurrence; (III) transforming predictors into disease probabilities or scores; and (IV) converting these into disease occurrences. Each step offers several alternatives for modeling the overall dependence structure. We apply our framework to a case study of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. By employing five combinations of methodological alternatives, we evaluate the approaches’ ability to achieve target characteristics and demonstrate their specific strengths and weaknesses. Matching the data-generating process with the estimation method allows for the successful recovery of input information, such as coefficients and correlations. Target properties like disease prevalence and associations are achieved to varying degrees depending on the methods used. While the proposed theory-driven framework is broadly applicable beyond the specific medical use case, it relies on careful, domain-informed parameter curation to generate meaningful synthetic datasets. Its flexible, adjustable input settings enable researchers to tailor data generation to their precise methodological requirements, providing a controlled basis for simulation-based validation without implying direct clinical inference.