Skip to content
Open access

Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

Abstract

Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

Read PDF

Similar papers

Jul 2026

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.

A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al. · 0 citations
Review Open access Jul 2026

Large Language Models in Spine Surgery: A Clinical Decision- Making Framework for the Next Decade with Emphasis on Degenerative Spine Care and LMIC Applications

A narrative review examines the evolving role of artificial intelligence in spine surgery, with particular emphasis on large language models in degenerative conditions of the cervical and lumbar spine, and proposes a structured six-level clinical decision-making framework spanning initial patient contact to postoperative care.

S. Ganesh, Jeena Joseph · 0 citations
Open access Aug 2026

Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study

The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.

Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al. · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations
#small language model Open access Aug 2026

Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis

Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.

Marco di Maio, G. Stopper, Vincenzo Di Matteo et al. · 0 citations