Skip to content
Open access

Guideline concordance of large language models for ERAS colorectal surgery recommendations: a blinded, clinician-rated comparison of Google Gemini and ChatGPT

Jun 2026 · Cukurova Anestezi ve Cerrahi Bilimler Dergisi · 0 citations · 9 references

TL;DR

Both LLMs showed high overall concordance with ERAS colorectal recommendations, however, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.

Abstract

Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;

Read PDF

Similar papers

Review Open access Aug 2026

Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts

LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.

Hetal Lad, Emily S Kwon, Ayushi Chadha et al. · 0 citations
Jul 2026

Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal Adenocarcinoma.

ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.

Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al. · 0 citations
Review Aug 2026

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al. · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations