Skip to content
Review Open access

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Jul 2026 · BMC Medical Informatics and Decision Making · 0 citations

Abstract

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited. This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated. Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds. Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery. Not applicable.

Read PDF

Similar papers

Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

INTRODUCTION Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO. METHODS A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores. RESULTS ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores. CONCLUSIONS There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Aug 2026

Quality and readability of AI Chatbot responses to frequently asked questions from patients undergoing progressive collapsing foot deformity surgery: a comparative study of ChatGPT, Perplexity, and Gemini.

BACKGROUND Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used? HYPOTHESIS The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials. PATIENTS AND METHODS Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall's W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons. RESULTS All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity. DISCUSSION AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to our hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education. LEVEL OF EVIDENCE III; cross-sectional comparative study.

L. Micicoi, J. Brué, Saumith Menon et al. · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

OBJECTIVE To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high-quality guidance is critical for safety. STUDY DESIGN Cross-sectional content analysis. SETTING Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics-based distribution. METHODS Twelve frequently asked questions about tracheostomy care were identified using search-listening tools and clinician input, then submitted to 5 AI chatbots - ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 - and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics. RESULTS Gemini 2.0 achieved significantly higher completeness scores than physician responses (P < .001), with DeepSeek and Grok 3 (P < .05) also outperforming (P < .05). Accuracy did not differ significantly between AI- and expert-generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring (P < .05). All AI and expert responses exceeded the NIH-recommended 6th-grade reading level, ranging from 10th-13th grade (P < .001). Inter-rater reliability was 78%. CONCLUSION AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI-generated responses. Future efforts should focus on adapting AI-generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations
Open access Aug 2026

Comparing the readability of AI-generated and society-authored patient information leaflets in orthopaedics

Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics. A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at p  < 0.05. Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories ( p  < 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2–10.6] vs. 8.7 [7.8–9.6], p  < 0.01). FRE was consistently lower in AI-generated texts across all categories ( p  < 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all p  < 0.01), indicating greater reading complexity. AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.

N. Kharma, S. Gill, Chayan Shanmugaratnam et al. · 0 citations
Open access Feb 2026

ChatGPT as a source of surgical information: Evaluation of responses to patient questions on hallux rigidus fusion

Background Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery. Methods Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch–Kincaid Grade Level, Gunning Fog, Coleman–Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC). Results ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68–0.80). Conclusion ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.

Kamil Balaban, M. Ertan, Mahmut Kalem · 0 citations
Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

SUMMARY OBJECTIVE: Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults. METHODS: Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]). RESULTS: Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3–11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64–0.80). CONCLUSION: While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults’ health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations