Skip to content

Interpretable Daily Inflow Forecasting Model via Cross-Attention Hybrid Neural Networks Integrated with SHAP Analysis

Jul 2026 · Water resources management · Vol 40 · 0 citations · 40 references

TL;DR

The case study for one-day-ahead forecasting indicates that the proposed framework achieves overall superior predictive performance relative to the benchmark methods according to six evaluation metrics and the Wilcoxon signed-rank test result, confirming the effectiveness of the proposed cross-attention fusion strategy and SHAP-guided input optimization.

View source

Similar papers

Open access Jul 2026

HoST: integrating Heuristic knowledge with attention-based LSTM networks for Thunderstorm prediction.

The proposed framework integrates attention-enhanced recurrent modelling with physically informed heuristic constraints with physically informed heuristic constraints, allowing the model to capture complex nonlinear atmospheric dynamics while maintaining meteorological consistency.

Kalyan Chatterjee, Mudassir Khan, Bhoomeshwar Bala et al. · 0 citations
Oct 2026

Enhancing Daily Streamflow Prediction: A Hybrid Deep Learning Framework Integrating Temporal Convolutions, Kolmogorov–Arnold Networks, and Multihead Attention

Accurate and reliable daily streamflow prediction is essential for effective water resources management, reservoir operation, drought mitigation, and early flood warning, yet it remains a persistent challenge due to the highly nonlinear, multiscale, and nonstationary characteristics of hydrological systems. To address these complexities, this study proposes DTCN-KAN-MA, a novel hybrid deep learning framework that synergistically integrates three key components. Dynamic dilated temporal convolutional networks (DTCNs) hierarchically extract short- to long-term temporal patterns across multiple scales. Kolmogorov–Arnold networks (KANs), replace traditional fixed activation functions with learnable spline-based functions to explicitly capture intricate nonlinear rainfall-runoff transformation. A multihead attention (MA) mechanism adaptively emphasizes the most informative hydrometeorological variables at each time step. The model was evaluated using comprehensive, basin-specific data sets from two hydroclimatically contrasting regions—the semi-arid Lanzhou station on the Yellow River and the humid Hekou station on the Diaojiang River. Results showed that DTCN-KAN-MA achieved Nash–Sutcliffe efficiency (NSE) values of 0.98 at both sites, comparable to the best individual benchmark (TCN), with a difference of only 1–2 percentage points. While the improvement in variance explanation was marginal, the combined model demonstrated lower error metrics, yielding RMSE–observations standard deviation ratio (RSR) scores of 0.14 and 0.12. Thus, the hybrid approach offers refined accuracy over standalone methods rather than a substantial gain in overall fit. Robustness analyses under input noise and ablation experiments further confirmed the model’s stability and the pivotal role of the KAN module in boosting generalization. The results indicated that the DTCN-KAN-MA model combined data-driven forecasting capabilities with improved predictive accuracy in the two studied basins, suggesting its potential utility for operational streamflow prediction in similar climatic settings. The model generated accurate daily streamflow forecasts for the two studied basins, offering a viable tool for informing flood warning, reservoir operation, and irrigation management.

Yaru Wu, Dong Wang, Qingwen Deng et al. · 0 citations
Jul 2026

Improving Daily Streamflow Forecasting under Non-stationarity with a Physics-Informed, Decomposition-Enhanced Deep Learning Model

Accurate and reliable streamflow forecasting is essential for hydropower generation, aquatic ecosystem management, and integrated water resource use. However, with ongoing climate change, streamflow series are increasingly exhibiting pronounced non-stationarity and non-linearity, which poses significant challenges for traditional forecasting models. Traditional physical models, despite having clear hydrological mechanisms, have limited accuracy under complex conditions. In contrast, deep learning models offer strong predictive capabilities but lack physical interpretability. Consequently, existing single-modeling approaches struggle to balance predictive accuracy with physical plausibility and show significant deficiencies in simulating extreme hydrological events. To address these issues, this study developed a multi-stage hybrid modeling framework integrating physical mechanisms, signal decomposition, and self-attention deep learning. This framework combines the Hydrologiska Byråns Vattenbalansavdelning (HBV) conceptual hydrological model, Variational Mode Decomposition (VMD), and a self-attention Bidirectional Long Short-Term Memory network (att-BiLSTM) to systematically enhance predictive performance for complex hydrological processes. Applied to the upper Heihe River Basin, the developed HBV-VMD-att-BiLSTM hybrid model demonstrates excellent testing set performance. It achieved Nash-Sutcliffe Efficiency (NSE) and Kling-Gupta Efficiency (KGE) values of 0.978 and 0.985, improvements of 29.0% and 15.5% over the standalone HBV model. Regarding extreme flow simulation, the hybrid model markedly reduced prediction biases for high flows (FHV improved from −13.4% to −1.7%) and low flows (FLV reduced from 9.2% to 2.6%), while maintaining high stability and robustness across different hydrological seasons. These findings provide critical technical support for refined water resource management and risk mitigation of extreme hydrological events in catchments facing significant hydrological non-stationarity.

Yunchao Jiang, Linsong Wang, Boliang Cui et al. · 0 citations
Open access Aug 2026

BiasCast: learning and adjusting real time biases from meteorological forecasts to enhance runoff predictions

The findings highlight the value of training strategies that allow models to directly learn bias correction during forecast transitions, emphasize the operational potential of combining sequential processing with near real-time discharge observations and identify physiographic catchment characteristics as key modulators of forecast skill improvement across diverse hydroclimatic settings.

O. Konold, Moritz Feigl, Patrick Podest et al. · 2 citations
Open access 2026

Explainable Spatiotemporal Ensemble Modeling Using Knowledge Distillation to Forecast Food Security

Accurate prediction of food security is crucial for managing climate-induced agricultural hazards, but current approaches struggle to include spatiotemporal complexity while producing interpretable and understandable results. This study fills this gap by conducting a thorough examination of machine learning paradigms for rainfall-driven food security forecasting in Uganda, including the use of baseline models, ensemble approaches, and spatiotemporal graph neural networks. A unifying framework was established, incorporating 15 various models, including linear regressors (Ridge and Lasso), tree-based ensembles (XGBoost and Gradient Boost), and a unique Spatial Graph Attention Network (GAT) with multi-head temporal attention. Temporal-GNN was implemented to model spatiotemporal trends in rainfall-food security interactions by combining graph convolutional networks (GCNs) and gated recurrent units (GRUs). To bridge the transparency gap in AI-driven predictions, SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) are systematically applied across all models, quantifying the influence of climatic drivers such as immediate rainfall intensity (r1h), seasonal fluctuations, and pressure gradients (rfq). Key advancements in this study include: XGBoost achieving state-of-the-art performance ( $\mathrm{R}^{2}=0.9949$ , MAE = 1.4067), surpassing baseline models by 14.3% in predictive accuracy; The Spatial GAT model demonstrating robust temporal dependency capture (MSE = 0.2092, MAE = 0.337), offering granular spatiotemporal insights despite a lower R2 (0.7801) and model-agnostic explainability revealing divergent feature importance. Linear models prioritise temperature, whereas ensembles emphasise atmospheric pressure dynamics. Recognising the practical challenges of deploying complex models, this work successfully implemented Knowledge Distillation (KD) to create a lightweight, efficient model from the best-performing XGBoost. The distilled student model achieved a compression ratio of 37.5% (reducing from 80 to 50 estimators) while maintaining 98.7% of the teacher model’s R2 score (0.9866 vs 0.9949). Despite a moderate increase in error metrics (MAE increased from 1.4067 to 2.3001, MSE from 3.8404 to 10.0075), the student model preserved the same feature importance hierarchy, with r1h and r1h_avg remaining the most influential predictors post-distillation. A reproducible pipeline incorporating spatial cross-validation, CUDA-accelerated training, and interactive XAI visualisation is introduced to enhance methodological rigour. Policymakers benefit from actionable trade-offs: the Voting Regressor (MAE = 2.2880) balances interpretability and performance, while the Spatial GAT enables localised, climate-resilient planning. This work advances scalable food security analytics by unifying statistical robustness with temporal dependency modelling, setting the stage for hybrid ensemble-GAT architectures in future research.

Ronald Atuhaire, Edward Kaboggoza, George Ssemaganda et al. · 0 citations