Skip to content
Open access

Accuracy and limitations of artificial intelligence chatbots in answering patient questions on scoliosis surgery.

Sep 2026 · Cirugía y Cirujanos · 0 citations
Medicine

Abstract

Objective The aim of this study was to analyze the accuracy, reliability, and quality of the content of responses provided by artificial intelligence (AI)-based chatbots to frequently asked questions related to scoliosis surgery. Methods A set of 25 questions related to diagnosis, treatment options, surgical risks, and recovery was developed based on clinical experience. The responses were anonymized and evaluated by two orthopedic surgeons using a 4-point scale to rate accuracy, clarity, and consistency. Cohen's kappa coefficient was used for interrater reliability, and the Kruskal-Wallis test was used for statistical analysis. Results ChatGPT 4o performed best, with 60% of responses rated as excellent. Gemini Advanced 1.5 Pro followed with 56%, while Microsoft Copilot performed poorly, with 76% of responses requiring moderate clarification. The differences between the chatbots were statistically significant (p < 0.001). ChatGPT 4o and Gemini did not differ significantly from each other (p = 0.850), but both outperformed Copilot (p < 0.001). Conclusions AI chatbots show great potential for patient education in scoliosis surgery. ChatGPT 4o was the most effective, although complex queries still require medical supervision.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.