Skip to content
Review

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Aug 2026 · Clinical and Translational Oncology · 0 citations · 15 references
Medicine

TL;DR

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

View source

Similar papers

Jul 2026

Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal Adenocarcinoma.

ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.

Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al. · 0 citations
Open access Aug 2026

Guideline Concordance of Large Language Models in the Management of Ureteral Stones: A Clinical Vignette-Based Comparative Study

ObjectiveLarge language models (LLMs) are increasingly used to answer medical questions, but their reliability in guideline-based urological decisions remains uncertain. This study aimed to evaluate the concordance of three widely used LLMs with the European Association of Urology (EAU) 2026 guideline recommendations for the active management of ureteral stones.Materials and MethodsIn this cross-sectional, vignette-based comparative study, the EAU 2026 ureteral-stone treatment algorithm was converted into 40 standardized clinical vignettes (four groups of ten: proximal 10 mm, distal 10 mm). ChatGPT, Gemini, and Claude were queried with the same standardized prompt, which did not name a specific guideline. Responses were scored against predefined EAU-based reference answers using a binary system. Concordance was compared with Cochran’s Q test.ResultsA total of 120 LLM-generated responses were evaluated. Overall concordance was 96.7% (116/120). ChatGPT achieved complete concordance (40/40, 100%), while Gemini and Claude each achieved 95% (38/40); the difference was not significant (Cochran’s Q=2.67, p=0.264). Concordance was complete in all 10 mm scenarios requiring URS prioritization. All four discordances were “incorrect prioritization” in >10 mm stones, presenting shock-wave lithotripsy as co-equal to ureteroscopy; each involved cross-guideline conflation with American Urological Association (AUA) framing. No unsafe recommendation or guideline hallucination was observed under the predefined scoring categories.ConclusionThe evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones. Concordance was complete in non–priority-sensitive

Unknown authors · 0 citations
Open access Jul 2026

The doctors of the future: the competition of ChatGPT-4, ChatGPT-4 omni, and Gemini 2.0 Flash in andrology.

OBJECTIVES Large language models (LLMs) are increasingly being used in medical research and as clinical decision-support tools. This study aimed to compare the accuracy and reliability of responses generated by large language models in response to andrology-related questions. MATERIALS AND METHODS Seventy questions concerning diagnosis, treatment, and general information were developed on the basis of the 2024 Andrology Guidelines of the European Association of Urology (EAU). These questions were submitted to three large language models, namely ChatGPT-4, ChatGPT-4o, and Google Gemini 2.0 Flash. The responses were independently evaluated by three senior urologists using a four-point rating scale. A total score (TS) > 9 indicated a good response, 6 ≤ TS ≤ 9 indicated a moderate response, and TS < 6 indicated a poor response. In addition, the self-correction capabilities of the models were evaluated, and changes in response accuracy after re-evaluation were analyzed. RESULTS ChatGPT-4o achieved the highest total scores in the diagnosis and treatment categories (p < 0.001). Google Gemini 2.0 Flash generated the longest responses but demonstrated the lowest accuracy. ChatGPT-4o also showed the greatest improvement following the self-correction process (Cohen's d = - 1.214, p < 0.01). Fleiss' kappa coefficient values ranged from 0.61 to 0.80, indicating substantial interrater agreement among the urologists. CONCLUSION ChatGPT-4o emerged as the most reliable model for andrology-related questions, providing responses that are were highly consistent with current clinical guidelines. The self-correction capabilities of the models improved response accuracy, suggesting that error-awareness mechanisms in large language models have the potential for further refinement. Nevertheless, expert supervision remains essential for the safe implementation of AI-assisted systems in clinical practice.

U. Uysal, E. Alma, A. Altunkol et al. · 0 citations
Review Aug 2026

Accuracy, Completeness, and Clarity of an AI-Based Chatbot for the EAU Neuro-Urology Guidelines.

This study aimed to externally validate the performance of the European Association of Urology (EAU) Guidelines Bot in neuro-urology by assessing the accuracy, completeness, and clarity of chatbot-generated answers to guideline-based questions and to compare its performance with that of a general-purpose large language model (ChatGPT 5.5). A cross-sectional validation study was conducted using 47 questions derived from the EAU Neuro-Urology Guidelines. Each question was linked to a specific recommendation and classified by recommendation strength (strong vs weak). Questions were independently submitted to both the EAU Guidelines Bot and ChatGPT 5.5 without additional prompting. Two expert urologists independently evaluated each response for accuracy, completeness, and clarity using a five-point Likert scale; discrepancies were resolved by a third reviewer. Overall, 45 questions (95.7%) were linked to strong recommendations and two (4.3%) to weak recommendations. The EAU Guidelines Bot and ChatGPT 5.5 achieved identical mean accuracy scores (4.96 ± 0.20), with all responses rated as highly accurate (Likert 4-5). ChatGPT 5.5 indicated significantly higher completeness scores than did the EAU Guidelines Bot (4.74 ± 0.44 vs 4.57 ± 0.54; p = 0.011), whereas clarity scores were not significantly different (4.83 ± 0.38 vs 4.77 ± 0.43; p = 0.083). High-quality completeness was observed in 46/47 EAU Guidelines Bot responses (97.9%) and 47/47 ChatGPT responses (100%). Score discrepancies between systems were identified in ten of 47 questions (21.3%) and were limited to completeness and clarity domains. Performance remained uniformly high across recommendation grades, with no meaningful differences observed. The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions. Its performance was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses. Although ChatGPT 5.5 generated more comprehensive answers, the EAU Guidelines Bot maintained closer adherence to the original guideline recommendations. Although not a substitute for clinical judgment, the tool appears to be a reliable adjunct for rapid access to evidence-based neuro-urological guidance.

Sabrina De Cillis, Riccardo Lombardo, Daniele Amparore et al. · 0 citations
Review Open access Aug 2026

Evaluating the potential of ChatGPT as an educational decision-support tool for hemodialysis decision-making in nephrology training

Background This study aimed to evaluate the performance of ChatGPT in identifying hemodialysis (HD) indications from authentic nephrology consultation notes and to compare its recommendations with both real-world clinical decisions and expert nephrologist consensus. Methods This exploratory observational study included 22 anonymized nephrology consultation notes from routine inpatient care at a tertiary care university hospital. Each note was independently evaluated by ChatGPT 5.4 using a standardized zero-shot prompt. The same notes were independently reviewed by three blinded senior academic nephrologists. The majority-vote consensus among these nephrologists was defined as the primary reference standard. Real-world decisions documented by nephrology fellows were evaluated as a secondary comparator. Agreement was assessed using Cohen’s kappa coefficient, and inter-rater reliability among nephrologists was evaluated using Fleiss’ kappa. Results Expert consensus classified eight of 22 cases (40.9%) as requiring hemodialysis and 13 (59.1%) as not requiring HD. Inter-rater agreement among the nephrologists was excellent (Fleiss’ κ = 0.814, p < 0.001). ChatGPT and real-world clinical decisions agreed in 18 of 22 cases (81.8%; κ = 0.633, p = 0.003), whereas ChatGPT and expert consensus agreed in 15 of 22 cases (68.2%; κ = 0.374, p = 0.069). Expert consensus and real-world clinical decisions agreed in 17 of 22 cases (77.3%; κ = 0.553, p = 0.007). Among the nine expert-defined HD cases, ChatGPT agreed in seven cases and demonstrated exact indication-level agreement in four cases. Among the 10 cases classified as HD by both ChatGPT and real-world clinical decisions, exact indication-level agreement was observed in four cases. Discrepancies were most commonly observed in cases involving metabolic acidosis, volume overload, and oliguria/anuria. Conclusion These exploratory findings suggest that ChatGPT may align more closely with nephrology fellows’ decision-making than with senior nephrologists’ judgments. The observed discrepancies may reflect differences in how clinical findings were interpreted and incorporated into the overall clinical context. Although ChatGPT cannot replace expert clinical judgment, it may have potential as an educational tool for trainees, encouraging systematic evaluation of HD indications and structured clinical reasoning. Larger, prospective, multicenter studies are needed to confirm these findings and evaluate ChatGPT’s educational impact in nephrology training.

Esin Avsar Kucukkurt, Feyza Bora, H. Sozel et al. · 0 citations