Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Toward Trustworthy AI Software Evaluation: A Controlled Benchmark of Deep Learning Architectures for 24-h Photovoltaic Power Forecasting

Accurate 24 h photovoltaic (PV) power forecasting is essential for day-ahead scheduling, storage operation, reserve planning, and market participation. However, published deep learning comparisons are often difficult to reproduce and interpret because they use inconsistent datasets, forecasting horizons, baselines, evaluation metrics, and leakage-control procedures. From a software engineering perspective, this limits the trustworthiness, comparability, and practical adoption of AI-based forecasting systems. This paper presents a controlled and reproducible benchmarking framework for evaluating AI-driven forecasting software. The framework is applied to nine deep learning architectures, three non-deep learning reference models, and two persistence baselines for hourly PV-power forecasting at a 350 kWp rooftop installation near Edinburgh, Scotland. All models were evaluated under a consistent experimental protocol, including the same chronological train–validation–test split, a 32-feature meteorological and solar-geometry input set, a 24-step forecasting horizon, capacity-normalised mean absolute error (NMAE), and Bayesian hyperparameter optimisation. The results show that TCN-LSTM achieved the best aggregate H24 performance with 7.22% NMAE, narrowly outperforming CPWformer-DEC at 7.28% and CT-PatchTST at 7.31%. LightGBM ranked fourth at 7.35% with fixed hyperparameters, outperforming six of the nine deep learning models. The top three models differed by only 0.09 percentage points, indicating that architectural superiority cannot be established reliably without significance testing and operational diagnostics. Per-horizon analysis showed that CT-PatchTST and S-Mamba performed best at the nearest forecast steps, whereas TCN-LSTM provided the most stable far-horizon profile. Peak-power diagnostics further revealed that aggregate NMAE can mask operational shortcomings, as Naive Persistence outperformed all deep learning models in high-output peak detection. The findings highlight the importance of reproducible benchmarking, leakage safeguards, horizon-aware evaluation, and operationally meaningful diagnostics in trustworthy AI software evaluation. The novelty of this work lies not in proposing a new architecture but in a controlled, reproducible framework that benchmarks fourteen forecasters under identical conditions, with explicit leakage safeguards, per-horizon reporting, and operationally meaningful peak diagnostics, enabling claims of architectural superiority to be made trustworthy rather than merely favourable. Architecture selection for PV forecasting should therefore consider not only aggregate accuracy but also reliability, interpretability of evaluation outcomes, and deployment-relevant performance behaviour.

Husein Mauladdawilah, Mohammed Balfaqih, Zain Balfagih et al. · 0 citations