Sep 2026· Langenbeck's archives of surgery (Print)· Vol 411· 0 citations· 37 references
Medicine
TL;DR
GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer, however, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integration into oncologic workflows.
Abstract
Large language models (LLMs) such as GPT-4 are being evaluated for their use as supportive tools in oncological treatment planning. However, in pancreatic cancer, current studies are confined to predefined question–answer formats, while studies specifically investigating real-world scenarios that benchmark LLM performance against multidisciplinary tumor board (MDT) decisions are lacking. This prospective comparative analysis evaluated treatment and diagnostic recommendations for patients with newly diagnosed or suspected pancreatic cancer between an MDT and GPT-4. Using MDT referrals, clinical data were entered into a clinical data matrix and submitted to GPT-4 for therapeutic and diagnostic recommendations. Outputs were assessed before and after additional prompting with 41 high-ranking abstracts relevant to pancreatic cancer care. The primary endpoint was the concordance of recommendations between the MDT and GPT-4 before and after literature-based prompting. Between September 1, 2024 and March 31, 2025, 45 patients were enrolled. The overall concordance rate between the MDT and GPT-4 was 73.3% (κ = 0.64, p < 0.0001) and did not improve following literature prompting. Discordance most often occurred in complex clinical scenarios. Concordance was highest in cases of metastatic disease (90.0%) and in neoadjuvant settings (90.0%) while it was lowest in patients requiring additional diagnostic workup (50.0%). GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer. However, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integration into oncologic workflows.
LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
Jun-Jo Sung, Eui Hyuk Chong, Incheon Kang et al.· Journal of Medical Internet...· 0 citations
While LLMs demonstrate promising concordance in standardized thyroid cancer management, they are best positioned as supportive decision aids—such as in MDT preparation and workflow streamlining—rather than replacements for expert multidisciplinary evaluation, particularly in complex clinical scenarios.
B. B. Büyük, Arzu Or Koca, F. Toprak et al.· Laryngoscope Investigative O...· 0 citations
Background/Objectives: Breast cancer management relies on multidisciplinary team (MDT) decisions that integrate clinical, radiological, pathological, and patient-related factors. Large language models (LLMs) may support such decisions, but evidence based on real-world cases remains limited. Methods: We evaluated the ag...
G. Dindelegan, Noé Yoshi François Poupel, George Ionuț Golea et al.· Journal of Clinical Medicine· 0 citations
ChatGPT achieved the highest overall concordance, although all models generated clinically acceptable recommendations in most cases, and all LLMs demonstrated high concordance with consultant-led thyroid cancer MDT decisions.
A. White, Kerry A. Leyton, N. Patel et al.· Updates in Surgery· 0 citations
The data demonstrate that the integration of LLMs in today’s MDT workflow is feasible and may benefit the quality of decision-making in specific cases, and suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making.
C. Buhr, L. Müller, D. Pinto dos Santos et al.· JMIR AI· 0 citations