Open Hydroclimatic Data and Iterative Machine Learning for River Discharge Forecasting in Babahoyo, Ecuador: Benchmarking, Uncertainty, and Operational Limits Without In Situ Validation
Abstract
This study evaluates the potential and limitations of open/model-derived hydroclimatic data for multi-horizon river discharge forecasting in Babahoyo, Ecuador. A quantitative, applied, retrospective longitudinal design used daily data for 2017–2026. The discharge target is the GloFAS v4 seamless product—reanalysis until July 2022, archived operational forecast thereafter; meteorological predictors are ERA5/ERA5-Land/IFS reanalysis, not forecasts. The framework combined leakage-aware feature engineering, temporal validation, rolling-origin backtesting, naïve baselines, machine-learning regression, conformal prediction intervals, and high-flow classification. Performance was strongest at one day, where models reproduced the signal closely (R2 = 0.909), although persistence remained highly competitive. Skill deteriorated at t + 7 and t + 14, where peak timing and magnitude became unreliable. Interval coverage was near-nominal at t + 1 but unreliable at longer horizons. The high-flow classifier identified most q90 cases, yet moderate precision and the absence of gauge validation prevent operational warning claims. Because the target is a 5 km grid simulation of a channel 100–150 m wide, the metrics quantify agreement with GloFAS, not with the physical river, and are reported to three significant digits. Overall, the study is a conservative benchmark for open hydroclimatic data in data-limited tropical floodplains: useful for exploratory monitoring and uncertainty diagnosis, but not a substitute for local hydrometric validation.