Forecasting Inflation with Machine Learning and Traditional Time-Series Models: A Multi-Horizon Rolling-Origin Assessment
Abstract
Introduction: The role of inflation forecasting in the monetary-policy assessment, financial planning and macroeconomic decision-making is crucial. This study also compares the Seasonal Autoregressive Integrated Moving Average (SARIMA) and Extreme Gradient Boosting (XGBoost) models to forecast Sticky Price Consumer Price Index (Sticky CPI) in the United States of America (USA) for seasonality during the past decade. Methodology: The current-vintage dataset (January 1968 to August 2024) contains 680 observations. All the 200 test months (Jan. 2008-Aug. 2024) were set aside for pseudo-out-of-sample evaluation, while the rest of the months (Jan. 1968-Dec. 2007) were used for model development only. One SARIMA order was selected one time from the pre-2008 sample and then these coefficients were re-estimated at each forecast origin. Recursive forecasts were obtained from a single SARIMA specification. XGBoost was fine-tuned separately for each horizon using five validation folds with a growing window that were completely nested within the pre-2008-time frame, with the hyperparameters fixed and the corresponding direct models re-estimated at each origin. Results: The reproduced analysis does not hold true for the 12-month XGBoost advantage. SARIMA records the lowest RMSE at all four horizons: 0.0869, 0.2420, 0.4229 and 0.9611, compared with XGBoost values of 0.1940, 0.4314, 0.7564 and 1.3040. After Holm adjustment, SARIMA is favored by Harvey-Leybourne-Newbold corrected Diebold-Mariano tests at 1, 3 and 6 months compared to XGBoost, while it is not significant at 12 month (p = 0.0673). Under a moving-block bootstrap, the 12-month RMSE is significantly larger for the predefined post-2020 analysis for all models, with SARIMA having a much smaller RMSE than XGBoost during that time. Additionally, feature ablation indicates that the entire nonlinear specification of XGBoost is weaker than a linear ridge model with the same features, after regularisation. Prediction intervals are conservative for both types of models, and SARIMA intervals are narrower and have lower interval scores. Conclusion: The results provide support to the discipline of benchmark evaluation, uncertainty reporting and monitoring of model-combination superiority through machine-learning, but do not support unconditional superiority of machine-learning.