Context-Aware Concept Distillation (CACD) is proposed, a framework developed in collaboration with domain experts to distill opaque LSTMs into interpretable, hydrology-aware surrogate models and a Residual Hypernetwork that dynamically modulates these concepts based on static basin characteristics.
Abstract
Effective flood risk management relies on accurate forecasting, yet the"black box"nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing Explainable AI (XAI) methods offer local attributions, they fail to provide the verifiable, operationally meaningful causal narratives required by disaster response authorities. To address this societal challenge, we propose Context-Aware Concept Distillation (CACD), a framework developed in collaboration with domain experts to distill opaque LSTMs into interpretable, hydrology-aware surrogate models. We introduce an unsupervised pipeline to discover a"Hydrological Language"and a Residual Hypernetwork that dynamically modulates these concepts based on static basin characteristics. Evaluated on 5,203 basins globally, our model achieves high fidelity (Median NSE 0.70), significantly outperforming black-box baselines (e.g., Multi Layer Perceptrons) on unseen future data. By demonstrating that human-interpretable concepts are sufficient to reconstruct flood dynamics, this work balances AI accuracy with the transparency required for responsible environmental decision-making.
Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have not explicitly represented the tacit expert rules, review checkpoints, and workflow constraints that connect model outputs to operational warning decisions. To address this issue, we propose HydroAgent, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning. We validated its effectiveness using five state-of-the-art LLMs in the South Yamhill River basin. Our results demonstrate that prior judgment captures observed peak flow and flood volume within 5% tolerance in 10 and 11 out of 14 events, with 5-fold cross-validation over 129 events yielding Pearson correlations of 0.62 and 0.84. Building on a high-baseline scheme library (average KGE 0.890), the guided scheme selection further improves KGE by 0.023-0.154, with simulated peak flow and flood volume falling within the prior judgment ranges for 14 and 13 out of 14 events. All five tested LLMs successfully execute the HydroAgent workflow with comparable judgment accuracy (40%-80%), while showing moderate performance variation and substantial cost differences. HydroAgent does not aim to replace human forecasters; instead, it translates their tacit expertise into an auditable and reproducible workflow, streamlining analytical steps and supporting more informed decision-making. This skill-orchestrated paradigm demonstrates how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.
Qingyi Yang, S. Qiu, Bingyao Li et al.· 0 citations
Wildfires are among the most severe natural hazards, posing a significant threat to both humans and natural ecosystems. The growing risk of wildfires increases the demand for forecasting models that are not only accurate but also reliable. deep learning (DL) has shown promise in predicting wildfire danger; however, its adoption is hindered by concerns over the reliability of its predictions, some of which stem from the lack of uncertainty quantification. To address this challenge, we present an uncertainty-aware DL framework that jointly captures epistemic (model) and aleatoric (data) uncertainty to enhance short-term wildfire danger forecasting. In the next-day forecasting, our best-performing model improves the area under the precision-recall curve by 1.1% and reduces the expected calibration error by 1.5% compared to a deterministic baseline, enhancing both predictive skill and calibration. Our experiments confirm the reliability of the uncertainty estimates and illustrate their practical utility for decision support, including the identification of uncertainty thresholds for rejecting low-confidence predictions and the generation of well-calibrated wildfire danger maps with accompanying uncertainty layers. Extending the forecast horizon up to ten days, we observe that aleatoric uncertainty increases with time, showing greater variability in environmental conditions, while epistemic uncertainty remains stable. Finally, we show that although the two uncertainty types may be redundant in low-uncertainty cases, they provide complementary insights under more challenging conditions, underscoring the value of their joint modeling for robust wildfire danger prediction. In summary, our approach significantly improves the accuracy and reliability of wildfire danger forecasting, advancing the development of trustworthy wildfire DL systems.
Spyros Kondylatos, N. Papadopoulos, G. Camps-Valls et al.· Machine Learning: Earth· 5 citations· ⚡1
Climate risk assessments for critical infrastructure are essential to identifying and predicting vulnerabilities early in the asset life cycle, enabling proactive mitigation through the implementation of technical and nature-based solutions (NbS) before impacts occur. However, such assessments often rely on dense quantitative indices that are difficult for non-technical stakeholders to interpret. To address this challenge, this paper presents an open-source decision support platform that combines OpenStreetMap site characterization, qualitative pre-screening, a quantitative IPCC AR6-aligned risk chain, and a downstream NbS recommendation layer. The approach deploys Large Language Models (LLMs) to translate analytical outputs into accessible narrative explanations. End-to-end site-characterization processing across three European demonstration sites took between 29 and 70 s. An exploratory ablation study investigated the faithfulness of the AI-generated explanations using three complementary metrics, demonstrating that the generated hazard assessments remained factually grounded and free from fabricated numerical values. Introducing example reports (exemplars) into the prompt context further stabilized the reliability of the output for complex risk indicators. Finally, a small blind expert evaluation with six researchers from adjacent technical domains provided convergent evidence: five of six raters independently rated with-exemplar Hazard Reports higher on completeness; among the five raters who expressed a directional preference, all five favored the with-exemplar condition (sign test, p = 0.031). Furthermore, seven of eight aggregate dimension-level comparisons confirmed that with-exemplar reports scored at least as high as their ablated counterparts.
Farid Arabameri, J. Ploennigs, M. Imani et al.· Infrastructures· 0 citations
The rapid proliferation of sensor networks connected infrastructure, and data-driven governance has positioned machine learning (ML) as a central instrument for managing modern urban environments. From adaptive traffic signal control to predictive policing, energy load forecasting, and public health surveillance, smart city platforms increasingly rely on opaque, high-capacity models such as deep neural networks and gradient-boosted ensembles to generate decisions that materially affect citizens’ lives. This dependence on black-box predictors has, however, exposed a fundamental tension between predictive performance and decision accountability, prompting a growing body of research into explainable machine learning (XML) as a mechanism for restoring transparency, auditability, and public trust. This review synthesizes the state of the art in explainable machine learning as applied to smart city decision-making, organizing the literature along four axes: the taxonomy of explainability techniques (intrinsic versus post-hoc, model-specific versus model-agnostic, and local versus global explanations); the core mathematical and algorithmic foundations underpinning dominant methods, including SHAP, LIME, counterfactual explanations, and attention-based interpretability; the architectural patterns through which explainability is embedded into smart city pipelines spanning edge, fog, and cloud tiers; and the domain-specific applications in transportation, energy, public safety, healthcare, and environmental monitoring. A comparative analysis of representative techniques is presented against criteria of fidelity, computational cost, stability, and human interpretability, supported by a consolidated case study on explainable traffic incident prediction. The review further surveys the software ecosystem supporting explainable smart city analytics, benchmark datasets, and evaluation metrics for explanation quality, before critically examining unresolved challenges including the fidelity-interpretability trade-off, adversarial manipulation of explanations, scalability under streaming data, and the absence of standardized regulatory frameworks. Ethical and societal dimensions, including algorithmic fairness, data governance, and the risk of explanation-washing, are discussed in relation to emerging legislation such as the EU AI Act. The review concludes by outlining research directions toward causally grounded, human-centered, and regulation-aligned explainable AI systems capable of sustaining trustworthy algorithmic governance in future smart cities.
Mohini Mittal, Jyoti Duhan Rathee· International Journal of Res...· 0 citations
Climate change, longer drought periods, and increased human activities have led to forests being burnt with greater frequency and intensity, with the consequences being a risk to ecosystems and to biodiversity as well as to public personnel and their health. Current monitoring systems using drones or satellites have enhanced wildfire monitoring, but there is still limited integration of features across modalities, insufficient cybersecurity resilience, and limited capacity to dynamically predict wildfire risk. To overcome these challenges, this paper introduces an Adaptive Trust-Aware Hierarchical Multimodal Learning Network (ATHML-Net) that combines drone image features, satellite observations, meteorological factors, and terrain data in a single cyber-resilient learning system. The proposed model aims to quantify the credibility of the heterogeneous data sources, design a physics-informed wildfire propagation learning mechanism to learn from fire and its propagation process for accurate spatiotemporal risk prediction, and propose a mechanism to estimate feature reliability based on trustworthiness to fuse features from different sources, which is robust against cheating attacks. The proposed framework has been shown to be effective through extensive experiments on publicly available multimodal wildfire datasets. The ATHML-Net detection achieves 99.21% accuracy, with precisions and recalls of 98.96% and 98.84%, respectively, yielding an F1 score of 98.90% and an AUC of 99.47%, reducing wildfire risk prediction error to an RMSE of 0.081. The framework maintained an average inference time of 23.8 ms/frame across all cases and maintained a detection accuracy of 96.83% in adversarial attack scenarios, representing an average 3.8% improvement over the best baseline framework. These results are highly effective for real-time forest fire detection and proactive wildfire risk prediction, providing a robust, accurate, and cyber-resilient solution.
Santhosh Chandra Konduru, Sujatha Lakshmi Narra, Hareshbhai Miyani et al.· JOIV: International Journal...· 0 citations
State-of-the-art multivariate time-series forecasters can model complex temporal and cross-variable dependencies, yet their opaque representations provide limited insight into why a particular forecast is produced. This lack of transparency restricts their use in settings where practitioners must understand and assess the factors underlying a prediction. We introduce ConceptTS, an interpretable forecasting framework that organizes its predictions around named, human-readable concepts. ConceptTS uses a large language model to propose task-relevant concepts and generate executable labeling rules, translating the language model's domain knowledge into direct supervision without costly manual concept annotation. The proposed concepts are organized into three complementary bottlenecks that describe the historical context, local forecast intervals, and the full forecast horizon. A shared decoder combines representations derived from their predicted activations to construct the forecast, making the model's decision process explicit and supporting direct concept-level interventions. Experiments on the Beijing Multi-Site Air Quality dataset show that ConceptTS achieves accuracy competitive with strong black-box baselines while producing semantically meaningful concept activations.