Large language models’ performance in answering common patient questions about colonoscopy: an expert-based evaluation
Abstract
This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and communication quality, as assessed by experts using a rubric-based method. This comparative, expert-based observational study analyzed the responses of both models to 11 colonoscopy-related questions written in Turkish, simulating real patient inquiries. To ensure consistency and fairness, questions were asked in separate sessions using a standardized zero-shot approach without prior examples or context. The gastroenterology experts from various institutions in Türkiye were recruited using snowball sampling and evaluated the responses using a rubric-based scale across four domains. Data were collected online. A total of 40 gastroenterologists participated (mean age 39.1 ± 6.4 years). For ChatGPT-4, more than 70% of responses were rated ≥ 4 across all domains, with median overall index scores ranging from 4.1 to 4.7. In contrast, Google Bard demonstrated lower performance, with median scores ranging from 3.6 to 4.0 and lower proportions of high ratings (≥ 4), particularly in comprehensiveness (55.7%) and applicability (58.4%). Comparative analysis showed that ChatGPT-4 significantly outperformed Google Bard across all domains: accuracy (p = 0.003, effect size = 0.47), applicability (p < 0.001, effect size = 0.64), comprehensiveness (p = 0.004, effect size = 0.46), and communication (p = 0.001, effect size = 0.52). No significant differences were observed across evaluator subgroups based on academic title, prior ChatGPT use, or age (all p-values > 0.05). ChatGPT-4 outperformed Google Bard across all domains, providing more accurate, comprehensive, and patient-centered responses. Despite these findings, human oversight remains essential to ensure safety and reliability in clinical communication. It should also be noted that different versions of LLMs may yield different results.