Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Evaluation and comparison of large language model responses to patient questions after diagnosis of high-risk human papillomavirus infection: an expert-rated digital patient education study

Background After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation. Methods In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm. Results ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios. Conclusion Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.

Zhen Hao, Lin Wang, Yue Wu et al. · 0 citations