Jul 2026· British Journal of Surgery· Vol 113· 0 citations
TL;DR
OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs.
Abstract
Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons.
Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125.
OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases.
Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.
While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.
Y. Ergen, S. Teke, E. G. Başaran et al.· Journal of Pediatric Gastroe...· 0 citations
Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
M. Hamdan, A. Harati, A. Al-bakheet et al.· medRxiv· 0 citations
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
METHODS
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
RESULTS
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
CONCLUSIONS
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
BACKGROUND
Large language models (LLMs) have emerged as powerful transformer-based systems capable of capturing long-range dependencies and complex semantic relationships in clinical language. In this review paper, we first examine the technical foundations of medical LLMs, including transformer architecture, attention mechanisms, training paradigms, and retrieval-augmented generation.
RESULTS
We then survey documented applications in spine surgery and spinal care, highlighting moderate guideline concordance (46-67%) for diagnostic support, automated generation of operative notes and discharge summaries for administrative workflows, LLM-assisted literature review and manuscript drafting for research support (with ~68% novelty accuracy), and translation of complex surgical concepts into patient-friendly materials at a seventh-grade reading level. We next explore emerging multimodal models that integrate text, imaging, laboratory, and genomic data via cross-modal attention, demonstrating superior performance in holistic diagnostic and prognostic tasks.
DISCUSSION
Finally, we discuss key implementation challenges, including model accuracy and hallucinations; computational, privacy, and regulatory constraints under HIPAA/GDPR; and bias mitigation, to outline strategies for safe, effective, and equitable deployment.
CONCLUSION
By mapping technical capabilities to clinical and research use cases, this review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.
Fabio Galbusera, Andrea Cina· European spine journal· 0 citations