Skip to content
Open access

Machine Learning-Based Crop Yield Prediction Using FAO Global Agricultural Data: A Comparative Study of Linear Regression, Random Forest, and XGBoost

Aug 2026 · International Journal For Multidisciplinary Research · 0 citations · 28 references

Abstract

Accurate crop yield prediction is essential for food security planning, agricultural policy-making, and sustainable farming management. This study develops and compares three machine learning models Linear Regression, Random Forest, and XGBoost for predicting crop yield using the FAO (Food and Agriculture Organization) global agricultural dataset. The dataset spans 212 countries and regions, covers 10 major crop types, and encompasses the period 1961–2016, comprising 56,717 records. Following a systematic preprocessing pipeline including missing value verification, duplicate removal, and label encoding, all models were trained and evaluated under identical experimental conditions using an 80:20 train-test split (45,373 training / 11,344 testing records). Performance was assessed using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and the coefficient of determination (R²). Random Forest achieved the best predictive performance with R² = 0.9586, MAE = 6,724.19 hg/ha, and RMSE = 13,829.60 hg/ha. XGBoost demonstrated competitive performance with R² = 0.8115, while Linear Regression showed limited effectiveness with R² = 0.0608, confirming its inability to model the nonlinear dynamics of global agricultural data. Feature importance analysis identified Item Code (crop type) as the most influential predictor, followed by geographic region and temporal factors. These findings confirm that ensemble learning methods, particularly Random Forest, are well-suited for modelling complex agricultural yield patterns at global scale.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.