Skip to content

Data-efficient machine learning for photovoltaic power prediction: Explainability, uncertainty, and learning saturation

Sep 2026 · Journal of Transportation and Sustainable Technologies · 0 citations

Abstract

Reliable photovoltaic power prediction is increasingly important for integrating variable solar generation into modern electricity systems, yet many operational installations have only limited historical data. This study developed a data-efficient, explainable, and uncertainty-aware machine-learning framework using 412 daylight observations from a grid-connected photovoltaic plant in India. Ambient temperature, module temperature, irradiation, day of year, and cyclic time variables were used to predict AC power with Gaussian Process Regression (GPR), Support Vector Regression (SVR), and Extra Trees Regression (ETR). Pearson correlation and principal component analysis were used to examine predictor structure, while chronological validation, permutation importance, SHAP analysis, predictive intervals, and learning-saturation experiments were applied to evaluate robustness. Irradiation showed the strongest correlation with AC power (r = 0.97), while the first three principal components explained 96.6% of predictor variance. ETR achieved the best independent-test performance (R² = 0.997, RMSE = 22.080 kW, MAE = 16.402 kW, sMAPE = 2.496%), closely followed by SVR (R² = 0.996). SHapley Additive exPlanations analysis identified irradiation as the dominant predictor, with a mean absolute contribution of 253.384 kW. ETR retained R² = 0.995 using only 50 training observations, demonstrating strong data efficiency. The results show that reliable and interpretable photovoltaic prediction can be achieved from very limited field measurements when model selection and validation are carefully designed under data-scarce operating conditions.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.