Skip to content

Comparative Performance and Utility of Large Language Models in Generating Psychological Screening Checklists for Bariatric Surgery

Aug 2026 · Bariatric Surgical Practice and Patient Care · 0 citations · 11 references

Abstract

Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation. Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey post hoc tests to compare expert ratings. Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness, F (2, 6) = 11.5, p = 0.008, and specificity, F (2, 6) = 9.8, p = 0.013. LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.

View source