Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.
Aug 2026· European Journal of Surgical Oncology· Vol 52 10, pp.
112053
· 0 citations· 24 references
Medicine
TL;DR
Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.
Abstract
INTRODUCTION
This study evaluated concordance between treatment-pathway outputs generated from structured clinical text by the large language model (LLM) GPT-5.4 Thinking and multidisciplinary team (MDT) decisions for mid- and low rectal cancer.
Materials And Methods
This single-center retrospective study included 260 patients who underwent standardized assessment, MDT discussion, and curative-intent surgery between January 2022 and December 2023. After database lock, structured de-identified clinical text was entered into GPT-5.4 Thinking using prespecified preoperative and postoperative templates, with MDT decisions treated as real-world reference decisions. The primary endpoint was preoperative concordance; secondary endpoints included postoperative and overall concordance. Concordance metrics, directional discordance, baseline comparators, repeatability in a 50-case subset, and exploratory logistic regression models were assessed.
Results
Preoperative, postoperative, and overall concordance rates were 57.7%, 66.9%, and 46.9%, respectively; Cohen's κ values were 0.269 and 0.449 for the preoperative and postoperative stages. Preoperatively, directional discordance relative to MDT decisions included 48 potential under-intensification, 28 potential over-intensification, and 34 directionally indeterminate or heterogeneous cases. Compared with the majority-class baseline, the LLM had the same preoperative crude concordance but higher κ and balanced category-specific concordance; postoperatively, it exceeded the majority-class baseline across these metrics. Non-identical mapped categories across three runs occurred in 12/50 preoperative and 8/50 postoperative assessments despite identical inputs.
Conclusion
GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited. Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs.
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
M. Hamdan, A. Harati, A. Al-bakheet et al.· medRxiv· 0 citations
For select patients with locally advanced rectal cancer (LARC) who achieve a clinical complete response after neoadjuvant therapy, non-operative management (NOM) offers an alternative to surgery but introduces uncertainty and an intensive surveillance burden. Decisions between NOM and surgery are preference-sensitive and evolve as treatment response becomes clearer. How clinicians operationalize shared decision-making (SDM) for NOM remains poorly understood. We aimed to characterize clinician communication strategies and decision processes surrounding NOM, focusing on risk framing and management of uncertainty. Eighteen colorectal surgeons, medical oncologists, and radiation oncologists from two academic medical centers completed 30–60-min semi-structured interviews. Transcripts were coded inductively and analyzed using constant comparison methods and iterative codebook development. Clinicians described discussions about NOM as a longitudinal, multidisciplinary process that unfolds across multiple encounters and clinical specialties, with medical oncologists and surgeons contributing at different points along the treatment trajectory. Complex information was organized through evidence-based explanations and various risk-framing approaches, allowing patients to engage in iterative deliberation while setting clear expectations about the conditional nature of NOM and the demands of ongoing surveillance. The timing and emphasis of discussions varied by treatment phase and clinical role, which supported gradual patient understanding but also created the potential for unrealistic expectations when recommendations from the treatment team were not clearly aligned. Decision-making about NOM for LARC unfolds over time under clinical uncertainty and across specialties. Structured, phase-specific communication and improved interprofessional alignment may strengthen shared decision-making in this preference-sensitive context.
E. Alagoz, Diana Gutierrez-Meza, Ana C. De Roo· Supportive Care in Cancer· 0 citations
Gastrointestinal stromal tumors (GISTs) are molecularly heterogeneous neoplasms whose management depends on individualized, multidisciplinary decision-making. While multidisciplinary tumor boards (MTBs) represent the standard of care, access remains limited in many clinical settings. This study evaluates the performance of two large language models in generating GIST MTB recommendations and assesses their agreement with expert MTB decisions using predefined clinical evaluation criteria. This retrospective single-center study included 99 GIST cases discussed at an institutional MTB. A structured prompt was developed to extract clinical variables and generate treatment recommendations. ChatGPT-5 and Qwen3 were independently evaluated across five predefined domains: diagnostic recommendations, therapeutic modalities, treatment sequence and timing, systemic therapy regimen selection, and clinical contextualization. Two expert reviewers scored all outputs in a blinded fashion. Normalized scores, inter-model comparisons, perfect-case rates, and inter-rater agreement were analyzed. Both models demonstrated high concordance with expert MTB recommendations, with mean total normalized scores of 0.901 for ChatGPT-5 and 0.875 for Qwen3, without a significant difference between models (p > 0.05). Perfect agreement was observed in 52.5% of ChatGPT-5 cases and 48.5% of Qwen3 cases (p > 0.05). Diagnostic recommendations scored significantly lower than all other domains in both models (all adjusted p < 0.05). Overall inter-rater agreement was almost perfect (weighted Cohen’s kappa=0.978). Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them. Diagnostic reasoning represented the weakest domain, reflecting the challenge of reconstructing context-dependent workup decisions from tumor board documentation. These findings support a potential assistive role for LLMs in GIST MTB workflows, while underscoring the continued necessity of expert oversight.
ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.
Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al.· Journal of Surgical Oncology· 0 citations