Guideline concordance of large language models for ERAS colorectal surgery recommendations: a blinded, clinician-rated comparison of Google Gemini and ChatGPT
Jun 2026· Cukurova Anestezi ve Cerrahi Bilimler Dergisi· 0 citations· 9 references
TL;DR
Both LLMs showed high overall concordance with ERAS colorectal recommendations, however, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.
Abstract
Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.
Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al.· Journal of Surgical Oncology· 0 citations
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.
Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al.· Clinical and Translational O...· 0 citations
LLM optimisation improves patient preference for trial information compared with standard registry descriptors, supporting further evaluation of its use in rendering patient-facing materials.
M. Tran, Kate Saw, Jeremy Mo et al.· Journal of Clinical Oncology· 0 citations
Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.
Yuanze Wei, Yulong Tian, Xiaodong Liu et al.· European Journal of Surgical...· 0 citations