Skip to content
Open access

Artificial Intelligence Chatbots as Sources of Cancer Pain Information: A Comparative Evaluation of Quality, Transparency, and Readability

Sep 2026 · Journal of Pain Research · Vol 19 · 0 citations · 44 references
Medicine

Abstract

Purpose To compare the quality, transparency, educational value, and readability of cancer pain information generated by four AI chatbots and to assess whether the outputs met prespecified readability benchmarks for patient education. Methods This online cross-sectional comparative study was conducted on July 22, 2026. Nine unmodified Google-Trends-derived cancer-pain-related queries were submitted to ChatGPT 5.5, Microsoft Copilot, Google Gemini 3.5 Flash, and Perplexity, yielding 36 responses. Two oncology clinicians independently assessed the responses using DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six established indices assessed readability. Matched model comparisons used Friedman tests with query as the repeated unit and Kendall’s W as the omnibus effect size; significant quality outcomes were followed by Holm-adjusted paired Wilcoxon signed-rank tests. Results DISCERN did not differ significantly across models (χ2(3)=6.682, P=0.083, W=0.247). EQIP differed across models (χ2(3)=20.721, P<0.001, W=0.767), with Copilot scoring higher than ChatGPT, Gemini, and Perplexity (Holm-adjusted P=0.023 for each). JAMA also differed (χ2(3)=19.645, P<0.001, W=0.728); Copilot and Perplexity scored higher than Gemini (Holm-adjusted P=0.023 and.039, respectively). GQS differed overall (χ2(3)=10.500, P=0.015, W=0.389), although no pairwise comparison remained significant after Holm adjustment. All six readability indices differed overall across models after Holm correction; descriptively, Gemini showed the greatest estimated reading difficulty. Model medians exceeded the prespecified grade-level benchmarks, and FRES medians were below 80. Conclusion The four chatbots differed across information quality, source transparency, educational utility, and formula-based readability. Claim-level clinical accuracy and safety were not evaluated in this study and warrant separate guideline-based assessment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.