Evaluation and comparison of large language model responses to patient questions after diagnosis of high-risk human papillomavirus infection: an expert-rated digital patient education study
Background After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation. Methods In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm. Results ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios. Conclusion Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.