Semantically-Guided Hydro-Synthesis for Mitigating Severe Inundation Visual Paucity in Vision-Based Urban Flood Monitoring Using Latent Diffusion Models
Abstract
The increasing unpredictability of weather patterns due to climate change necessitates robust, real-time flood monitoring systems for high-risk disaster-prone cities. Existing monitoring approaches primarily utilize satellite-based remote sensing, in-situ telemetry such as ultrasonic level detection and precipitation gauges, or convolutional neural networks trained on heterogeneous, crowd-sourced internet imagery. However, the effectiveness of these systems is limited by Severe Inundation Visual Paucity (SIVP), which refers to the scarcity of historical visual data depicting extreme, high-stage inundation events required for training deep learning classifiers. To address this data gap, this study presents Semantically-Guided Hydro-Synthesis (SGHS), a novel generative framework that employs latent diffusion models to synthesize photorealistic, weather-variant inundation overlays on fixed CCTV perspectives. Unlike prior generative flood-data approaches that operate from satellite or aerial perspectives, produce numerical depth grids, or rely on site-agnostic image priors, SGHS is uniquely conditioned on a site-specific UAV-reconstructed 3D digital twin, employs prompt-driven domain randomization across six atmospheric parameters, and is governed by a formalized ten-level anthropometric severity taxonomy that enables graduated, decision-relevant severity estimation rather than binary flood/no-flood inference. By using SGHS to generate a balanced dataset encompassing escalating flood severities and diverse meteorological conditions, the SIVP constraint is directly mitigated, thereby enriching the training distribution with essential high-stage instances absent from historical records. Empirical results are reported across three evaluation tiers with disaggregated test-set provenance and across ten random seeds with mean and standard deviation. On the Tier-1 sim-to-real holdout at Levels 0 through 3, the raw-data baseline achieves Accuracy <inline-formula> <tex-math notation="LaTeX">$=0.1990~\pm ~0.0146$ </tex-math></inline-formula> and F<inline-formula> <tex-math notation="LaTeX">$1=0.2170~\pm ~0.0160$ </tex-math></inline-formula>, while the SGHS-augmented curated classifier achieves Accuracy <inline-formula> <tex-math notation="LaTeX">$=0.7700~\pm ~0.0099$ </tex-math></inline-formula> and F<inline-formula> <tex-math notation="LaTeX">$1=0.7825~\pm ~0.0108$ </tex-math></inline-formula> (Wilcoxon signed-rank, p < 0.002). Matched-augmentation controls applying focal loss, balanced sampling, and standard photometric augmentation to the raw partition reach only F<inline-formula> <tex-math notation="LaTeX">$1=0.3120~\pm ~0.0211$ </tex-math></inline-formula>, isolating the contribution of generative augmentation. Tier-3 results at Levels 6 through 9 (F<inline-formula> <tex-math notation="LaTeX">$1=0.6685$ </tex-math></inline-formula>–0.7135) are explicitly reframed as evidence of internal consistency within the simulation domain rather than as a claim of sim-to-real operational transfer, since extreme-severity ground truth remains structurally absent from the observational record. These results indicate that site-specific, semantically-guided generative augmentation provides a measurable and statistically significant improvement over both raw-data and algorithmic class-imbalance baselines, addressing hydrological data scarcity for vision-based disaster response systems.