Explainable Machine Learning for Exploratory Residual Screening of Electricity-Consumption Disclosures in a Sample of Chinese Listed Firms
Abstract
Firm-level deviations in reported electricity consumption can help prioritize corporate energy disclosures for verification. This study presents an explainable machine-learning workflow that predicts log-transformed annual electricity consumption from harmonized resource-use variables, firm-size controls, a previously available disclosure lag, stock-code-prefix and year indicators, and missingness indicators, then screens group-held-out cross-fitted residuals. The final sample comprised 354 firm years from 180 Chinese listed firms. LightGBM achieved pooled out-of-fold log-scale R2 = 0.6206, RMSE = 2.1094, and MAE = 1.3809. At |z| > 2.0, 23 observations were screened (12 positive, 11 negative); eight crossed the Gaussian-reference BH-FDR q < 0.05 screening boundary, and four crossed the Bonferroni boundary. All eight Tier 1 cases ranked 1–8 under distribution-free absolute-residual ranking, although empirical BH and Bonferroni adjustment yielded no discoveries. A heteroscedasticity-adjusted ranking remained strongly associated with the primary ranking (Spearman ρ = 0.9117; Top-20 overlap = 14/20). The outputs are verification priorities rather than confirmatory anomaly labels and do not imply inefficiency, misreporting, or causality.