Safety, accuracy, empathic communication, information quality, and readability of five large language model interfaces answering public questions about interstitial cystitis/bladder pain syndrome
Patients increasingly use large language model (LLM) interfaces for health information, but their safety and quality for public questions about interstitial cystitis/bladder pain syndrome (IC/BPS) remain uncertain. This study evaluated the safety, accuracy, empathic communication, information quality, reliability, and readability of five publicly accessible LLM interfaces. In this CHART-guided cross-sectional comparative study, 58 public-facing IC/BPS questions were submitted once, in English, to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao using a standardized single-turn, zero-shot protocol. Three blinded senior urologists independently assessed safety, accuracy, empathic communication, DISCERN, EQIP, JAMA benchmark criteria, and Global Quality Score. Six readability indices were calculated. Inter-rater agreement was significant for all manually assessed metrics. Fleiss’ kappa for safety was 0.822, and ICC(2,1) values for other rater-assessed metrics ranged from 0.761 to 0.848. Unsafe responses occurred in all interfaces, ranging from 5.2% for ChatGPT to 8.6% for DeepSeek and Doubao, without a significant between-interface difference (Cochran Q = 1.000, p = 0.910). Accuracy and empathic communication differed significantly across interfaces (both p < 0.001). ChatGPT had the highest median accuracy score, whereas DeepSeek had the highest empathic communication score. Information-quality and reliability scores also differed significantly (all p < 0.001); ChatGPT achieved higher DISCERN, EQIP, and GQS scores, while Gemini achieved higher JAMA scores. Readability differed significantly across interfaces, but none met predefined patient-facing readability benchmarks. Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions. Unsafe responses were uncommon but present in all interfaces, and no interface consistently outperformed the others. LLM interfaces may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.