Machine Learning-Based Prediction of Six-Minute Walk Distance in Children with Obesity
Highlights What are the main findings? Explainable machine learning model, especially XGBoost, outperformed the conventional linear regression in predicting six-minute walk distance (6MWD), as it can better model complex and non-linear relationships between anthropometric parameters in children with obesity. Age is the main predictor of 6MWD in all models, with sex-specific different patterns of body mass index, height and waist circumference. What are the implications of the main finding? The use of SHAP-based explainable AI with machine learning offers can bypass the limitations of conventional linear models and provide a transparent and accurate approach for the assessment of functional exercise capacity in pediatric obesity. These newly derived ML-informed prediction equations are practical tools for field-based clinical assessment tailored to each subgroup, but external validation in large cohorts is required before routine implementation. Abstract Background: The six-minute walk test (6MWT) assesses functional exercise capacity, but existing reference equations for children with obesity rely on traditional linear regression, potentially overlooking complex, non-linear relationships between anthropometric characteristics and functional exercise capacity. Objective: This study aimed to develop and internally validate machine-learning (ML) prediction models and preliminary prediction equations derived from explainable ML models for six-minute walk distance (6MWD) in Tunisian school-aged children with obesity and to compare their predictive performance with a conventional regression-based approach. Methods: We analyzed data from 236 school-aged children with obesity (104 females, 132 males; 6–12 years). Anthropometric measurements included body mass (BM), height, body mass index (BMI), waist circumference (WC), and hip circumference (HC). Five models were evaluated: linear regression, Ridge, Lasso, Elastic Net, and XGBoost. Performance was assessed using five-fold cross-validation and evaluated by the mean absolute error (MAE) and root mean square error (RMSE). Predictor importance was assessed using SHapley Additive exPlanation (SHAP) and Gini importance. Results: XGBoost achieved the best predictive performance, with the lowest MAE (17.4 ± 2.5 m in females and 19.5 ± 1.9 m in males) and RMSE (25.3 ± 3.1 m in females and 26.4 ± 2.2 m in males). Age was the strongest predictor across all models (SHAP: 54.2–62.0%; Gini importance: 0.52–0.69), followed by height and BMI. Sex-specific analyses indicated that, in females, age and BMI contributed ~80% to the cumulative SHAP analysis; whereas, in males, age, height, and WC were the primary factors. Conclusions: ML, particularly XGBoost, significantly improves 6MWD prediction in school-aged children with obesity compared with traditional linear regression. Explainable ML increases model interpretability by evaluating the relative importance of anthropometric predictors. These obesity-specific prediction models may better capture complex non-linear associations between anthropometrics and functional exercise capacity. These initial population-specific models should be validated in larger independent samples before routine use clinical or field settings.