Skip to content

Author

Emrecan Akgün

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This study aimed to compare responses generated by four widely used LLMs. Methods: Seventeen frequently asked questions regarding PFPS were identified through Google searches and adapted into lay language. The questions were submitted to OpenAI GPT-5, Google Gemini 2.5 Pro, xAI Grok 4, and DeepSeek-V3.2-Exp using a standardized patient scenario. A total of 68 question-specific responses were independently evaluated by four orthopedic surgeons using the DISCERN instrument. Inter-rater reliability was assessed using the intraclass correlation coefficient. Readability was evaluated using the Gunning Fog Index, Coleman–Liau Index, and Flesch Reading Ease Score. Between-model comparisons were performed using the Friedman test, followed by Bonferroni-adjusted pairwise analyses. Results: The omnibus Friedman test showed a significant between-model difference in DISCERN scores (p = 0.002). In Bonferroni-adjusted pairwise comparisons, GPT-5 had lower DISCERN scores than Gemini 2.5 Pro (adjusted p = 0.006), Grok 4 (adjusted p = 0.021), and DeepSeek-V3.2-Exp (adjusted p = 0.036), whereas no significant differences were observed among the other three models. However, the absolute differences were small, and the between-model difference was not significant in the sensitivity analysis using the median evaluator score (p = 0.381). Inter-rater agreement was moderate for GPT-5 and DeepSeek-V3.2-Exp but poor for Gemini 2.5 Pro and Grok 4. Readability differed significantly among the models across all three indices. DeepSeek-V3.2-Exp generally showed more favorable numerical readability values, whereas Grok 4 tended to produce more difficult text; however, no model was consistently superior across all readability measures. The median Gunning Fog and Coleman–Liau scores for all four models exceeded the commonly recommended sixth- to eighth-grade reading level for patient education. Conclusions: The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability. Their responses may supplement general patient education, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making. LLM-generated information should be critically reviewed and should not replace assessment by a qualified healthcare professional.

Oktay Polat, Berk Koncalıoğlu, M. Gündoğdu et al. · 0 citations