Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
Abstract
Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.
A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al.· Skeletal Radiology· 0 citations
A narrative review examines the evolving role of artificial intelligence in spine surgery, with particular emphasis on large language models in degenerative conditions of the cervical and lumbar spine, and proposes a structured six-level clinical decision-making framework spanning initial patient contact to postoperative care.
S. Ganesh, Jeena Joseph· Nepal Journal of Neuroscienc...· 0 citations
The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.
Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al.· Journal of Medical Internet...· 0 citations
Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.
Yuanze Wei, Yulong Tian, Xiaodong Liu et al.· European Journal of Surgical...· 0 citations
Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.
Marco di Maio, G. Stopper, Vincenzo Di Matteo et al.· Bioengineering· 0 citations
OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs.
A. Ravishankar, C. Lam, A. Bulloso et al.· British Journal of Surgery· 0 citations