Skip to content
Review Open access

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Jul 2026 · British Journal of Surgery · Vol 113 · 0 citations

TL;DR

OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs.

Abstract

Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons. Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125. OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases. Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.

Read PDF

Similar papers

Jul 2026

The illusion of competence: Evaluating the clinical reasoning of large language models in pediatric gastroenterology.

While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.

Y. Ergen, S. Teke, E. G. Başaran et al. · 0 citations
Open access Aug 2026

Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

M. Hamdan, A. Harati, A. Al-bakheet et al. · 0 citations
Open access Jul 2026

A multidimensional benchmarking framework for large language models in oncologic decision making

A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.

M. Halıcı, Serkan Saltürk, Irem Sayin et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

INTRODUCTION Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO. METHODS A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores. RESULTS ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores. CONCLUSIONS There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Review Open access Aug 2026

Large language models in spine care and research.

BACKGROUND Large language models (LLMs) have emerged as powerful transformer-based systems capable of capturing long-range dependencies and complex semantic relationships in clinical language. In this review paper, we first examine the technical foundations of medical LLMs, including transformer architecture, attention mechanisms, training paradigms, and retrieval-augmented generation. RESULTS We then survey documented applications in spine surgery and spinal care, highlighting moderate guideline concordance (46-67%) for diagnostic support, automated generation of operative notes and discharge summaries for administrative workflows, LLM-assisted literature review and manuscript drafting for research support (with ~68% novelty accuracy), and translation of complex surgical concepts into patient-friendly materials at a seventh-grade reading level. We next explore emerging multimodal models that integrate text, imaging, laboratory, and genomic data via cross-modal attention, demonstrating superior performance in holistic diagnostic and prognostic tasks. DISCUSSION Finally, we discuss key implementation challenges, including model accuracy and hallucinations; computational, privacy, and regulatory constraints under HIPAA/GDPR; and bias mitigation, to outline strategies for safe, effective, and equitable deployment. CONCLUSION By mapping technical capabilities to clinical and research use cases, this review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.

Fabio Galbusera, Andrea Cina · 0 citations