Skip to content

Author

Zhenhua Zhao

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Safety, accuracy, empathic communication, information quality, and readability of five large language model interfaces answering public questions about interstitial cystitis/bladder pain syndrome

Patients increasingly use large language model (LLM) interfaces for health information, but their safety and quality for public questions about interstitial cystitis/bladder pain syndrome (IC/BPS) remain uncertain. This study evaluated the safety, accuracy, empathic communication, information quality, reliability, and readability of five publicly accessible LLM interfaces. In this CHART-guided cross-sectional comparative study, 58 public-facing IC/BPS questions were submitted once, in English, to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao using a standardized single-turn, zero-shot protocol. Three blinded senior urologists independently assessed safety, accuracy, empathic communication, DISCERN, EQIP, JAMA benchmark criteria, and Global Quality Score. Six readability indices were calculated. Inter-rater agreement was significant for all manually assessed metrics. Fleiss’ kappa for safety was 0.822, and ICC(2,1) values for other rater-assessed metrics ranged from 0.761 to 0.848. Unsafe responses occurred in all interfaces, ranging from 5.2% for ChatGPT to 8.6% for DeepSeek and Doubao, without a significant between-interface difference (Cochran Q  = 1.000, p  = 0.910). Accuracy and empathic communication differed significantly across interfaces (both p  < 0.001). ChatGPT had the highest median accuracy score, whereas DeepSeek had the highest empathic communication score. Information-quality and reliability scores also differed significantly (all p  < 0.001); ChatGPT achieved higher DISCERN, EQIP, and GQS scores, while Gemini achieved higher JAMA scores. Readability differed significantly across interfaces, but none met predefined patient-facing readability benchmarks. Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions. Unsafe responses were uncommon but present in all interfaces, and no interface consistently outperformed the others. LLM interfaces may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.

Jiang-Tao Zhu, Zhenhua Zhao, Song Li et al. · 0 citations