Data diversity vs. model complexity in the prediction of pediatric bipolar disorder: Evidence from academic and community clinical samples.
Abstract
Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed to enhance diagnostic reliability. We evaluated a clinical decision tool (nomogram), statistical methods (logistic regression, LASSO), machine learning (support vector machine, random forest, k-nearest neighbors, extreme gradient boosting), and deep learning (multilayer perceptron) for pediatric bipolar disorder prediction across two datasets collected in academic (N = 550) and community (N = 511) clinical settings. We compared three modeling strategies: cross-dataset validation, cross-dataset with interaction terms, and pooled-dataset. We assessed model performance using discrimination, calibration, and predictor importance ranking. In the baseline cross-dataset approach, all models showed good internal discrimination in the academic dataset, but external discrimination in the community dataset substantially declined. Interaction-enhanced models slightly improved internal discrimination but not external performance or calibration. Recalibration substantially improved cross-dataset calibration. Models trained on the pooled sample showed strong performance on held-out samples drawn from the heterogeneous pooled cohort, with good calibration for most models. Across models and training strategies, PGBI-10M was consistently identified as the most important predictor. Predictive models for pediatric bipolar disorder showed strong internal performance but limited cross-setting generalizability due to dataset shift and miscalibration. Within the present study, increasing model complexity did not improve external performance, whereas training on pooled data improved performance on held-out samples from the heterogeneous pooled cohort. These findings suggest that training-data diversity may provide greater practical benefit than increasing model complexity for developing robust psychiatric prediction models, underscoring the importance of open and collaborative datasets.