Recurrent Deep Learning Models for PM2.5 Time Series Forecasting in Northern Thailand
Abstract
Forecasting PM2.5 concentrations remains a challenging data science problem because of strong nonlinearity, temporal dependence, and pronounced nonstationarity associated with seasonal and episodic pollution events. This study presents a comparative evaluation of recurrent neural network architectures for one-day-ahead PM2.5 forecasting using daily observations from four provinces in northern Thailand with varying pollution dynamics. Standard RNN, LSTM, and GRU models were developed within a unified forecasting framework using historical PM2.5 concentrations and meteorological variables as predictors. Model performance was evaluated on an independent test dataset using R2, RMSE, and MAE, together with Diebold–Mariano tests based on forecast error series. The experimental results indicate that the GRU model provides more robust forecasting performance under conditions of strong volatility and nonstationary behaviour, while the LSTM and RNN models remain competitive in comparatively more stable environments. No single model consistently dominated across all provinces, indicating that forecasting effectiveness depends strongly on the temporal characteristics of local air pollution dynamics rather than on architectural complexity alone. While all models capture the overall seasonal structure of PM2.5 variation, extreme pollution peaks are systematically underestimated partly because MSE-based training biases predictions toward the central tendency of the data distribution. However, gated architectures show improved responsiveness to abrupt concentration changes relative to the standard RNN. These findings highlight the importance of matching recurrent architectural design to the temporal regime of the target environment.