Development of a machine learning model for diabetes risk prediction
Abstract
Diabetes mellitus imposes a growing burden on health systems, yet the prediagnostic period, when prevention is still possible, is poorly characterized by existing prediction tools. This independent study develops and evaluates an endto-end longitudinal diabetes-risk modeling pipeline using twelve years of annual health-check data from 6,892 Thai patients (2005–2016). The task is framed asrolling patient-year at-risk classification with explicit control over two temporal dimensions: prediction horizon N ∈ {1, 2, 3, 4, 5} years and historical input window M ∈ {1, 3, 5} years. The pipeline compares biostatistical longitudinal classifiers, tree-based machine-learning models, and survival models on a shared cohort under a common data-engineering and patient-splitting protocol. No single model family is optimal across all horizons: tree-based models (CatBoost, XGBoost) dominate or remain near the top at short horizons, while biostatistical models (Logistic Regression, Generalized Estimating Equation) lead at longer horizons. The longest available history window (M =5) was the most beneficial configuration for the main direct-prediction finalists. Temporal framing is itself a substantive design choice, as illustrated by the two-stage survival variant restricted to a one-year history window (M =1), which collapsed to a degenerate ranker. A separate intervention-safe track develops monotonic models that preserve clinically plausible score direction under tested favorable lifestyle scenarios; across all horizons and tested scenarios these models were directionally correct in every case, yetmatched or exceeded the pure-prediction winners at horizons one through four. The central contribution is a dual-track recommendation system: a horizonspecific pure-prediction leaderboard for passive patient-risk screening, and an intervention-safe scoring benchmark for patient-facing what-if simulation. Sharing the same data engineering, patient split, and outcome definition but serving different operational objectives, the two tracks reflect that ranking patients by future risk and patient-facing what-if scoring are related but not interchangeable tasks. A complementary thesis-only ablation shows that the predictive value of the calendar-time features is concentrated at the one-year horizon and negligible beyond three years, motivating a construct-validity alternative in the deployed API.