Robustness Evaluation of Building Electricity Load Forecasting Models Using Similarity-Controlled Fourier Surrogates
Abstract
Building electricity load forecasting models are usually selected on an unchanged historical test set, although rankings may shift when temporal organization changes. We evaluated five models across 60 non-residential buildings using seven Fourier-surrogate settings with the executed zero-clipping correction, ten realizations per setting, and target offsets corresponding to effective leads of 2 and 25 h. Aggregate Pearson similarity declined from 0.984 at S7 to 0.373 at S1, but only 14/60 buildings were strictly monotonic. HistGradientBoosting and LightGBM had numerically lower median reference CVRMSEs than Ridge without statistically detectable paired reference advantages. Under this Fourier-phase-randomization-plus-zero-clipping operator, Ridge showed a lower observed-similarity degradation AUC, with absolute-scale evidence strongest at 2 h and inconclusive across site clusters at 25 h. Winner changes exceeding 5% occurred in 52.1% and 40.6% of the events in the two information-fair comparisons. The framework complements fixed-test accuracy with controlled, operator-specific model-selection sensitivity screening rather than a causal simulation of deployment drift.