A Rainfall-Only Hybrid LSTM-Random Forest Framework for Data-Driven Urban Flood-Risk Decision Support in Jakarta under Severe Class Imbalance
Abstract
Purpose – Urban flood-risk screening in Jakarta requires prediction approaches that remain useful when complete hydrological and infrastructural flood data are not consistently available. This study evaluates a rainfall-based decision-support framework for next-day flood-risk detection using Long Short-Term Memory (LSTM), Random Forest (RF), and a hybrid LSTM-RF model. Design/methods/approach – Rainfall predictors were reconstructed from ERA5-Land reanalysis and transformed into lagged, rolling-window, and seasonal features. Official flood-occurrence labels were obtained from BNPB’s DIBI records. The supervised dataset comprised 4,134 daily observations (132 flood and 4,002 non-flood events) and was chronologically split into training, validation, and test sets. Findings – At the fixed 0.50 threshold, the LSTM model achieved the highest ROC-AUC of 0.735516 and recall of 0.740741, although its precision remained low at 0.054945 because of many false positives. RF and Hybrid LSTM-RF achieved high accuracy of 0.967391 but failed to detect positive flood cases in the hold-out test set. Research implications/limitations – These findings show that accuracy is insufficient for evaluating rare-event flood prediction and that rainfall-only models require careful threshold calibration, official-label expansion, and additional hydrological predictors before being interpreted as operational flood-warning tools Originality/value – Flood-occurrence labels were constructed from the official tabular Data Informasi Bencana Indonesia (DIBI) records managed by BNPB, so the supervised target represented officially recorded flood events rather than rainfall-threshold proxy labels. The study also evaluates and compares LSTM, RF, and Hybrid LSTM-RF models using rainfall-only predictors for next-day flood-risk detection under severe class imbalance.