Skip to content

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

Sep 2026 · Phlebology · pp. 2683555261490374 · 0 citations · 16 references
Medicine

TL;DR

LLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis, however, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for patient education.

Abstract

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered information require thorough evaluation, especially in rare diseases like lipedema. Therefore, this study aimed to evaluate the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to questions frequently asked by patients with lipedema.MethodsThis cross-sectional study assessed the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to 25 commonly asked lipedema-related questions. Each model was queried twice in separate sessions, and answers were evaluated by three independent experts using a four-point rating scale. To ensure the objectivity and consistency of expert evaluations, inter-rater agreement was assessed using Cohen's kappa coefficient.ResultsDeepSeek achieved the highest proportion of comprehensive and correct responses (72%), followed by Gemini (64%) and ChatGPT (56%). Accuracy varied across content categories, with notable limitations particularly in treatment, follow-up, and maintenance questions. Reproducibility analysis revealed that DeepSeek produced the most consistent responses across sessions, while ChatGPT and Gemini showed more variability, particularly in treatment and quality-of-life questions. Cohen's kappa values indicated high inter-rater agreement overall, with perfect agreement in some categories for ChatGPT and DeepSeek.ConclusionsLLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis. However, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for patient education.

View source

Similar papers

Open access Oct 2026

Large language models’ performance in answering common patient questions about colonoscopy: an expert-based evaluation

This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and co...

Neslihan Güneş Aydemi̇r, D. Yapar, Yasemin Demi̇r Avcı et al. · 0 citations
#large language models Review Open access Sep 2026

Comparative performance of ChatGPT-4.0, DeepSeek, Gemini and Perplexity in answering common questions from patients with COPD

This study compared responses from the LLMs ChatGPT-4.0, DeepSeek, Gemini, and Perplexity to 20 English-language questions about COPD that patients could present and found no LLM consistently met the approximate eighth-grade target.

Zeng-Li Chen, Yun-Lan Jiang, Hui Jiang et al. · 0 citations
Open access Aug 2026

Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.

S. Wegmann, T. Rosenkranz, Philipp Egenolf et al. · 0 citations
Open access Sep 2026

A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate for selected basic pharmaceutical questions, and significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate.

Y. L. T. Bayala, I. A. Tinni, Thierry Boris Wend-Yam Yaméogo et al. · 0 citations
Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.