Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Evaluating machine learning and baseline methods for crop yield forecasting using small datasets

Accurate crop-yield forecasting is difficult when only small official-statistics datasets are available. This study evaluates a reproducible pre-season regional forecasting workflow for wheat, maize, and sunflower in Poltava region, Ukraine, using harmonised AgroStats records for 2010–2024. The analysis is based on one annual regional time series per crop; it is therefore a benchmark of temporal forecasting from official statistics rather than a gridded yield-prediction or spatial-mapping study. Forecast accuracy is assessed with mean absolute error (MAE), root mean squared error, and mean absolute percentage error. To avoid information leakage, feature transformations, imputations, scaling, and hyperparameter choices are performed only within the corresponding historical training window under a forward temporal design: training set 2010–2018, validation set 2019–2021, and held-out test set 2022–2024. In addition to ElasticNet, XGBoost, and LightGBM, the study compares transparent baseline forecasting methods: Naive, FORECAST.LINEAR, LINEST, and autoregressive integrated moving averag. Under the conservative lag-only scenario, the best 2022–2024 test results are obtained for maize (LightGBM, MAE 0.69 t ha−1) and sunflower (LightGBM, MAE 0.04 t ha−1), whereas for wheat the linear-trend baseline remains slightly better (MAE 0.49 t ha−1 versus 0.54 t ha−1 for ElasticNet). Supplementary analyses show that extended lag structures can improve selected crops and that seasonal NASA Prediction of Worldwide Energy Resources climate aggregates improve maize forecasts (MAE 0.52 t ha−1) but not wheat or sunflower. SHapley Additive exPlanations are used descriptively to examine whether the selected models rely on agronomically plausible predictors. The findings should be interpreted as crop-specific evidence under a small annual dataset: machine learning does not guarantee superiority over simple baselines, but it can provide a reproducible comparison framework and useful gains for selected regional forecasting tasks.

O. Kopishynska, Mark Fedorchenko, Y. Utkin et al. · 0 citations