Review
Open access
Jul 2026
184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery
OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs.
A. Ravishankar, C. Lam, A. Bulloso et al.
· British Journal of Surgery · 0 citations