Sep 2026· Journal of Medical Internet Research· Vol 28, pp. e101137-e101137· 0 citations· 24 references
Medicine
TL;DR
LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
Abstract
Abstract Background Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision environments remains unclear. In particular, the stability of LLM-generated recommendations—that is, whether a model produces the same answer when given the same clinical input—has rarely been examined. Objective This study aimed to evaluate the stability of treatment recommendations generated by contemporary LLMs when identical HPB cases are queried repeatedly, and their concordance with the treatment decisions reached at an institutional MDT conference. Methods This retrospective study included consecutive cases discussed at a single-center HPB MDT conference between September 1, 2024, and August 31, 2025. Standardized clinical case summaries derived from preconference documentation were provided to 4 LLMs (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5) through their consumer web interfaces. Each model recommended a treatment among predefined MDT treatment options, and identical queries were repeated 4 times in separate sessions. Stability was quantified as the discordance rate relative to the initial response and, without privileging any single query, as the mean pairwise agreement and Fleiss κ across the 4 iterations. Concordance with MDT decisions was assessed using both the initial and modal responses, together with Cohen κ and class-wise F1-scores. Results A total of 107 MDT cases were analyzed. Stability differed significantly across models (P=.01). Gemini 3 Pro showed the lowest discordance rate (mean 12.8%, SD 2.3%) and the highest reference-free agreement (Fleiss κ=0.737), whereas GPT-4o showed the highest discordance rate (mean 30.2%, SD 6.5%) and the lowest agreement (Fleiss κ=0.430). Concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and the highest-performing model differed between the 2 definitions. Class-wise F1-score was consistently lower for surgery (0.400‐0.520) than for chemotherapy (0.621‐0.836). Complete discordance occurred in 17 of 107 (15.9%) cases and in none of the 31 anatomically unresectable cases (Fisher exact test, P=.003). Recurrent or on-treatment disease (adjusted odds ratio [OR] 5.40, 95% CI 1.65-17.68; P=.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03-26.78; P=.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54-69.38; P=.016) were independently associated with complete discordance. Conclusions LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology. Importantly, response stability varied substantially across models, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools. These findings suggest that LLMs may serve as a reasoning-support layer in MDT-like decision environments, but their response stability must be systematically characterized before clinical integration.
Background/Objectives: Breast cancer management relies on multidisciplinary team (MDT) decisions that integrate clinical, radiological, pathological, and patient-related factors. Large language models (LLMs) may support such decisions, but evidence based on real-world cases remains limited. Methods: We evaluated the ag...
G. Dindelegan, Noé Yoshi François Poupel, George Ionuț Golea et al.· Journal of Clinical Medicine· 0 citations
GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer, however, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integrat...
F. Gehrisch, K. Kirkgöz, Antonie Willner et al.· Langenbeck's archives of sur...· 0 citations
While LLMs demonstrate promising concordance in standardized thyroid cancer management, they are best positioned as supportive decision aids—such as in MDT preparation and workflow streamlining—rather than replacements for expert multidisciplinary evaluation, particularly in complex clinical scenarios.
B. B. Büyük, Arzu Or Koca, F. Toprak et al.· Laryngoscope Investigative O...· 0 citations
The data demonstrate that the integration of LLMs in today’s MDT workflow is feasible and may benefit the quality of decision-making in specific cases, and suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making.
C. Buhr, L. Müller, D. Pinto dos Santos et al.· JMIR AI· 0 citations
In HCC management, AI demonstrated a high level of agreement with expert MDT decisions, suggesting its potential role as a complementary decision-support tool, however, limitations persist for elderly patients and borderline clinical scenarios, in which individualized human judgment remains essential.
C. Ciulli, Emanuele Scarpa, Pasquale Chiacchio et al.· Annals of Surgical Oncology· 0 citations
Interdisciplinary clinical boards (ICB) on chronic inflammatory diseases discuss complicated and multidisciplinary cases with inflammatory conditions among experts in their discipline. Medically specialized large language models (LLMs) are finding increasingly widespread use among physicians to answer clinical question...
M. Ronicke, L. Sollfrank, Ioannis Sagonas et al.· Scientific Reports· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.