Aug 2026· IEEE Access· Vol 14, pp. 133252-133263· 0 citations· 33 references
Computer Science
TL;DR
A comparative evaluation of six deep learning models--covering state-space, MLP, RNN, and Transformer-based architectures--emphasizing generalization across markets suggests that N-HiTS and NBEATSx perform competitively in limited-data scenarios, while transformer-based models can reach comparable accuracy but tend to require more adaptation and tuning.
Abstract
While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized comparison. As a result, many studies have relied on different datasets and metrics to evaluate methods in isolated settings, making it difficult to assess progress and compare state-of-the-art approaches consistently. In this work, we use public data to evaluate deep learning models for electricity price forecasting (EPF) across multiple market settings. Our goal is to establish a reproducible framework that enables a consistent evaluation of forecasting models. Although deep learning has been explored for day-ahead EPF, many prior studies are limited to single-market settings, narrow feature sets, or fixed training regimes. This work presents a comparative evaluation of six deep learning models--covering state-space, MLP, RNN, and Transformer-based architectures--emphasizing generalization across markets. We simulate low-data target-market conditions using zero-shot, one-shot, and few-shot learning. Our test set focuses on the Germany-Luxembourg (DE-LU) bidding zone in 2024 using a standardized dataset with calendar, historical price, and market-derived features. Our findings suggest that N-HiTS and NBEATSx perform competitively in limited-data scenarios, while transformer-based models can reach comparable accuracy but tend to require more adaptation and tuning. Model performance also benefits from careful feature selection and hyperparameter tuning, and we note that the differences between the strongest models are often small.
Forecasting commodity prices remains a challenging task due to market volatility, structural breaks, and changing economic conditions. This study evaluates the forecasting performance of classical econometric, deep learning, convolutional, and Transformer-based models for aluminum futures prices. Daily aluminum futures data are analysed using 11 forecasting approaches. Forecasting experiments are conducted for two distinct evaluation periods representing the years 2022 and 2025 to assess the robustness of model performance under different market environments. The empirical results reveal substantial differences in forecasting performance across model families. Recurrent neural network architectures, particularly GRU and RNN, consistently achieve the lowest forecasting errors across most horizons. The classical ARIMA model remains highly competitive despite its relative simplicity. In contrast, Transformer-based models generally fail to outperform simpler alternatives and frequently produce higher forecast errors. Statistical comparisons based on Diebold–Mariano tests and Model Confidence Set procedures indicate that performance differences among the best-performing models are often limited, suggesting that increased model complexity does not necessarily translate into superior predictive accuracy. Overall, the findings highlight the importance of empirical model evaluation and demonstrate that parsimonious forecasting approaches can remain effective competitors to substantially more complex machine learning architectures in aluminum futures price forecasting.
László Vancsura, Tibor Tatay, Tibor Bareith et al.· Decision Making Advances· 0 citations
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland, and Spain over 2021-2025. Their performance is assessed in terms of point and probabilistic forecasting accuracy, as well as economic value in battery energy storage arbitrage. Only the TabPFN models consistently and significantly outperform the benchmarks across all three markets and all statistical measures. However, this statistical dominance does not translate directly into economic dominance: TabPFN performs best under unlimited bids and riskier quantile-based strategies, whereas the Distributional Deep Neural Network benchmark is more profitable when risk tolerance is lower. Thus, foundation models cannot universally replace market-specific models, and their value depends on both model architecture and the decision problem.
Predicting electrical demand in distribution systems is a fundamental problem for the efficient operation of smart grids, especially under scenarios of high temporal variability. This study compares two multi-step forecasting strategies for 24-h horizons: a direct 24→24 strategy and a recursive strategy based on sequential 24→1 predictions. Four machine learning and deep learning architectures are evaluated: LSTM, N-HiTS, U-Net, and LightGBM, using real data from electrical feeders belonging to distribution systems in the equatorial region of Ecuador. The methodology includes constructing time windows, non-overlapping train/validation/test partitioning for evaluation, consistent normalization, and comparative analysis using MAE, RMSE, and MAPE metrics. The results show that the direct 24→24 strategy achieves the best overall performance, with LSTM standing out with an approximate MAPE of 4.12%. However, the recursive strategies exhibit greater stability in the face of atypical patterns observed during holidays and weekends. Furthermore, U-Net demonstrates competitive performance in both accuracy and temporal robustness, while LightGBM stands out for its computational efficiency. It is concluded that the selection of a forecasting strategy depends on the required balance between overall accuracy, temporal stability, and computational cost.
Erik Fernando Mendez-Garces, David Buldain, M. Comech· Energies· 0 citations
Short-term electricity demand forecasting is a critical enabler of the secure and efficient operation of modern power systems, particularly amid increasing renewable energy integration, smart grid expansion, and the broader energy transition. This paper presents a rigorous comparative analysis of electricity demand forecasting models, encompassing statistical methods, Machine Learning (ML), Deep Learning (DL), and hybrid architectures. A structured taxonomy is proposed to classify models according to their methodological family, application horizon, and data requirements, thereby providing a unified reference framework for researchers and energy-sector practitioners. Models are evaluated using a multi-criteria framework comprising accuracy, robustness, scalability, interpretability, computational cost, and the capacity to handle exogenous variables. The analysis identifies critical research gaps, including the limited integration of probabilistic forecasting into operational contexts and the absence of standardized evaluation protocols under real-world conditions. Future research directions are outlined, with particular emphasis on uncertainty quantification, adaptive learning strategies, and hierarchical forecast coherence in systems with high penetration of distributed energy resources.
A. Torres-Sánchez, Á. Jaramillo-Duque, W. Villa-Acevedo· Processes· 0 citations
Summary With the rapid integration of renewable energy and the ongoing advancement of electricity market, accurate real-time prediction of electricity clearing prices has become increasingly critical. This paper introduces a deep learning-based framework for predicting real-time clearing price intervals, which employs time-shift feature construction and a greedy-based feature selection process to optimize high-dimensional feature sets. To address the nonlinear and non-stationary nature of price data, signal decomposition techniques are integrated to improve the extraction of underlying feature patterns. A point prediction model is first developed using a long short-term memory (LSTM) network, after which the upper and lower bounds of the prediction interval are derived via a similarity-based matching algorithm and statistical interval analysis, thereby explicitly quantifying market uncertainty. Numerical experiments based on the Shanxi Province (China) electricity market dataset show that the proposed model achieves higher interval coverage compared to baseline intervals, while maintaining practically useful interval widths, which offering more reliable decision support for electricity market participants.
Jun Yang, Jian Zhou, Lei Zhu· iScience· 0 citations
Proper carbon price forecasting is the key to the dynamics of the emissions trading system (ETS) and assists market participants. The objectives of this study are to examine the price prediction of the carbon price, taking as the input previous price histories of major price ETS markets, but it is not a classification problem, but a time-series regression task. The aim of the forecasting is to forecast the value of future carbon prices at short-term scales with the help of supervised learning models. The structure of a hybrid deep learning network is suggested, whereby a Long Short-Term Memory (LSTM) network is trained with an Enhanced Pelican Optimization (EPO) algorithm to enhance hyperparameter selection. Model performance is measured in a rigorously chronologically rolling-origin validation scheme to prevent data leakage and may be more realistic for forecasting in the real world. Regular regression-based performance measures, such as MAE, RMSE, and MAPE, are used to guarantee that performance is comparable to the literature on carbon price forecasting. The experimental findings indicate that the proposed methodology yields consistent falls in error in comparison with the general statistical and deep learning baselines on the datasets analyzed. Although the results show that the methodology has been improved in terms of R2 in forecasts, the findings are limited to the chosen markets and forecast horizons. This makes the findings map conclusions to the predictions of performance and strength, and not a direct economic or policy influence. Further employment will involve more analysis of other markets, horizons, and economic utility analysis.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026