A Calibration-Aware Comparative Regression Framework for Local Rainfall Probability Estimation and Rain-Risk Classification
Abstract
Estimating rainfall likelihood supports agriculture, irrigation planning, and local weather decision-making. This study evaluates ten regression-based machine learning models for next-day rain occurrence using 346 daily weather records from Kuantan, Pahang, Malaysia, covering 2025. Weather and lagged features were used with linear, regularized, polynomial, kernel-based, and ensemble models. Raw regression outputs outside the valid probability range were clipped to 0–1, while model-specific rain/no-rain thresholds were selected using the chronological validation subset and fixed for independent testing. Elastic Net achieved the strongest decision performance, with an F1-score of 0.8533, accuracy of 0.7925, precision of 0.7805, and recall of 0.9412. LightGBM produced the lowest continuous-score error, with an RMSE of 0.4366 and Brier Score of 0.1906. MLR produced the highest invalid raw-output rate at 94.34%. The results show that probability-oriented and threshold-based metrics should be evaluated jointly when regression outputs support rainfall-risk decisions. The imbalanced validation subset remains a limitation and motivates longer multi-year and multi-station validation.