Skip to content

Author

Caner Kılıç

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

Large Language Models Versus Multidisciplinary Tumor Board Decisions in Thyroid Cancer

ABSTRACT Objectives Large language models (LLMs) are increasingly proposed as clinical decision‐support tools; however, their agreement with real‐world multidisciplinary tumor board (MDT) decisions remains insufficiently investigated in thyroid oncology. To evaluate the concordance between treatment recommendations generated by ChatGPT 5.2 and Gemini 3.0 and decisions made by a tertiary multidisciplinary thyroid tumor board. Methods This study included 59 consecutive patients discussed at a tertiary MDT between August and December 2025. Anonymized clinical data, including demographics, ultrasonographic findings, and Bethesda cytology, were provided to both LLMs using standardized structured prompts. MDT decisions were defined as the reference standard. Agreement was assessed using exact concordance rates and Cohen's kappa ( κ ) statistics with 95% confidence intervals. Results ChatGPT 5.2 achieved a concordance rate of 71.2% (42/59), demonstrating substantial agreement ( κ = 0.623; 95% CI 0.459–0.771). Gemini 3.0 showed a concordance rate of 64.4% (38/59), reflecting moderate agreement ( κ = 0.527; 95% CI 0.359–0.684). Discordance increased in complex scenarios involving lateral neck dissection, radioactive iodine therapy, and active surveillance. Conclusions While LLMs demonstrate promising concordance in standardized thyroid cancer management, they are best positioned as supportive decision aids—such as in MDT preparation and workflow streamlining—rather than replacements for expert multidisciplinary evaluation, particularly in complex clinical scenarios. Level of Evidence 3.

Barış Büyük Buyuk, Arzu Or Koca, Felat Toprak et al. · 0 citations