Skip to content

Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal Adenocarcinoma.

Jul 2026 · Journal of Surgical Oncology · 0 citations · 13 references
Medicine

TL;DR

ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.

Abstract

INTRODUCTION The management of rectal adenocarcinoma requires navigation of complex, branching guideline pathways encompassing neoadjuvant sequencing, surgical approach, organ preservation, and surveillance, yet real-world guideline adherence remains as low as 60-70%. The ability of current-generation large language models (LLMs) to accurately navigate these decision points has not been fully characterized.

Methods

In this cross-sectional, vignette-based study, 135 clinical questions were constructed from 45 pages of NCCN Rectal Cancer Guidelines (Version 4.2024). ChatGPT-4o was queried using standardized prompts with up to 3 clarifying questions permitted per query. Responses were independently evaluated by two physician raters on a 5-point Likert scale, with potential discrepancies adjudicated by a board-certified surgical oncologist. Primary outcomes were the proportion of responses rated Correct (score ≥ 3) and Accurate (score ≥ 4). Inter-rater reliability was assessed using Cohen's kappa, and subgroup analysis was performed across clinical domains using the Kruskal-Wallis test.

Results

Of 135 questions, 127 (94.1%; 95% CI, 88.7-97.0%) were Correct and 121 (89.6%; 95% CI, 83.3-93.7%) were Accurate. One hundred two responses (75.6%) were completely correct without additional prompting. Performance was consistent across clinical domains (Kruskal-Wallis H = 0.530, p = 0.767). Inter-rater agreement was perfect (κ = 1.0). Eight responses (5.9%) contained partially or wholly incorrect information, with errors concentrated in multi-step conditional treatment decision points.

Conclusion

ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by our group. The concentration of errors in complex conditional treatment algorithms suggests that LLMs excel at discrete factual recall but may struggle with multi-step reasoning under clinical uncertainty. Prospective validation using real-world clinical data and comparison with multidisciplinary tumor board recommendations remain necessary prior to clinical integration.

View source

Similar papers

Review Aug 2026

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al. · 0 citations
Open access Aug 2026

Guideline Concordance of Large Language Models in the Management of Ureteral Stones: A Clinical Vignette-Based Comparative Study

ObjectiveLarge language models (LLMs) are increasingly used to answer medical questions, but their reliability in guideline-based urological decisions remains uncertain. This study aimed to evaluate the concordance of three widely used LLMs with the European Association of Urology (EAU) 2026 guideline recommendations for the active management of ureteral stones.Materials and MethodsIn this cross-sectional, vignette-based comparative study, the EAU 2026 ureteral-stone treatment algorithm was converted into 40 standardized clinical vignettes (four groups of ten: proximal 10 mm, distal 10 mm). ChatGPT, Gemini, and Claude were queried with the same standardized prompt, which did not name a specific guideline. Responses were scored against predefined EAU-based reference answers using a binary system. Concordance was compared with Cochran’s Q test.ResultsA total of 120 LLM-generated responses were evaluated. Overall concordance was 96.7% (116/120). ChatGPT achieved complete concordance (40/40, 100%), while Gemini and Claude each achieved 95% (38/40); the difference was not significant (Cochran’s Q=2.67, p=0.264). Concordance was complete in all 10 mm scenarios requiring URS prioritization. All four discordances were “incorrect prioritization” in >10 mm stones, presenting shock-wave lithotripsy as co-equal to ureteroscopy; each involved cross-guideline conflation with American Urological Association (AUA) framing. No unsafe recommendation or guideline hallucination was observed under the predefined scoring categories.ConclusionThe evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones. Concordance was complete in non–priority-sensitive

Unknown authors · 0 citations
Open access Jul 2026

Performance of leading large language models in adhering to clinical guidelines for anaplastic thyroid cancer: a comparative study

Leading LLMs show variable capacity to align with ATC clinical guidelines, while top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use.

Mohamed Yasser, Ghada Barakat, S. Awny et al. · 0 citations
Review Aug 2026

Accuracy, Completeness, and Clarity of an AI-Based Chatbot for the EAU Neuro-Urology Guidelines.

This study aimed to externally validate the performance of the European Association of Urology (EAU) Guidelines Bot in neuro-urology by assessing the accuracy, completeness, and clarity of chatbot-generated answers to guideline-based questions and to compare its performance with that of a general-purpose large language model (ChatGPT 5.5). A cross-sectional validation study was conducted using 47 questions derived from the EAU Neuro-Urology Guidelines. Each question was linked to a specific recommendation and classified by recommendation strength (strong vs weak). Questions were independently submitted to both the EAU Guidelines Bot and ChatGPT 5.5 without additional prompting. Two expert urologists independently evaluated each response for accuracy, completeness, and clarity using a five-point Likert scale; discrepancies were resolved by a third reviewer. Overall, 45 questions (95.7%) were linked to strong recommendations and two (4.3%) to weak recommendations. The EAU Guidelines Bot and ChatGPT 5.5 achieved identical mean accuracy scores (4.96 ± 0.20), with all responses rated as highly accurate (Likert 4-5). ChatGPT 5.5 indicated significantly higher completeness scores than did the EAU Guidelines Bot (4.74 ± 0.44 vs 4.57 ± 0.54; p = 0.011), whereas clarity scores were not significantly different (4.83 ± 0.38 vs 4.77 ± 0.43; p = 0.083). High-quality completeness was observed in 46/47 EAU Guidelines Bot responses (97.9%) and 47/47 ChatGPT responses (100%). Score discrepancies between systems were identified in ten of 47 questions (21.3%) and were limited to completeness and clarity domains. Performance remained uniformly high across recommendation grades, with no meaningful differences observed. The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions. Its performance was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses. Although ChatGPT 5.5 generated more comprehensive answers, the EAU Guidelines Bot maintained closer adherence to the original guideline recommendations. Although not a substitute for clinical judgment, the tool appears to be a reliable adjunct for rapid access to evidence-based neuro-urological guidance.

Sabrina De Cillis, Riccardo Lombardo, Daniele Amparore et al. · 0 citations