Skip to content
Open access

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study

Aug 2026 · Frontiers in Public Health · 0 citations · 35 references

TL;DR

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation.

Abstract

Search-enabled large language model interfaces are increasingly used by the public for health information, but their performance in mpox-related public health consultation remains unclear. This study evaluated their safety, accuracy, empathy, reliability/information quality, and readability. We conducted a single-query comparative cross-sectional evaluation using 52 predefined mpox-related public consultation questions. Each question was submitted once to each of six search-enabled LLM interfaces, yielding 312 first responses. Responses were assessed against a guideline-based reference framework. Safety was coded as a binary outcome, while accuracy and empathy were rated on 5-point scales. Reliability/information quality was evaluated using DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was assessed using six established readability indices. Five trained raters independently evaluated the human-scored outcomes. Unsafe responses were relatively infrequent but occurred in all six interfaces, with safe-response rates ranging from 86.5 to 92.3%. No pairwise difference in Safety remained statistically significant after Benjamini–Hochberg correction. Overall differences across interfaces were statistically significant for Accuracy, Empathy, all four reliability/information quality measures, and all six readability indices. Benjamini–Hochberg-adjusted post hoc analyses identified outcome-specific pairwise differences, although the pairwise patterns varied across measures. The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability. Although unsafe responses were relatively uncommon, potentially harmful outputs occurred in every interface. These findings support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation. The results represent a time- and configuration-specific interface-level snapshot; they should not be attributed to the underlying base models in isolation or interpreted as establishing reproducible performance or a stable hierarchy across sessions, versions, or settings.

Read PDF

Similar papers

Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists.

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations
Open access Aug 2026

Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot

Background Patients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues. Methods This cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28–29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons. Results Interrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p < 0.001), and for GQS (p = 0.011). Within this defined late-May 2026 response set, responses returned by the interface displaying the label ‘Grok-4.3’ had the highest observed sample means for DISCERN (58.68), EQIP (81.14), GQS (4.27), and the JAMA visible metadata/transparency proxy score (1.00). No evaluated public-interface response set achieved recommended sixth-grade readability, and no individual response met all six readability thresholds. Responses returned by the interface displaying the label ‘DeepSeek-v4’ had the most favorable observed readability profile within this sampled response set, although the mean FKGL remained 11.53. These observations should not be interpreted as durable model rankings. Conclusion In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag. Pronounced ceiling and floor effects preclude conclusions about clinical sufficiency or safety. These findings should not be generalized to non-English use, different health-literacy levels, country-specific emergency-care pathways, other regions or account configurations, or later interface states.

Wei Zhong, Yuanyuan Zhang, Yu Huang et al. · 0 citations
Review Open access Jul 2026

Assessing multiple-choice question quality in internal medicine: a comparative analysis of three large language models against expert consensus

Background Large language models (LLMs) are increasingly explored for their potential to support quality assurance in medical education assessment. However, limited evidence exists on the alignment between LLM evaluations and expert judgment across multiple dimensions of multiple-choice question (MCQ) quality. Methods This comparative methodological study evaluated 85 MCQs from an internal medicine clerkship examination. Three LLMs (Claude Sonnet 4, Gemini 2.5 Flash, and Llama 3.3 70B Instruct Turbo) and three medical education experts independently assessed each question for cognitive level (Revised Bloom’s Taxonomy), alignment with the intended learning outcome (5-point Likert scale), and presence of technical flaws based on NBME guidelines. Agreement was calculated using Fleiss’ kappa for cognitive level classification and Cohen’s kappa for binary technical flaw criteria, with intraclass correlation coefficients (ICC) for Likert-scale alignment ratings. Results For cognitive level classification, Gemini (κ = 0.424, p < 0.001) and Claude (κ = 0.415, p < 0.001) showed moderate agreement with experts; Llama demonstrated lower agreement (κ = 0.266, p < 0.001). Alignment with learning outcomes yielded weak-to-moderate agreement for all models (ICC 0.121–0.380). For technical adequacy, Claude and Gemini achieved almost perfect agreement on detecting negatively worded stems (κ = 0.897, p < 0.001) and substantial agreement on inconsistent numerical data (κ = 0.661, p < 0.001) but showed poor agreement on more subjective flaws. Llama performed poorly across most technical criteria. Conclusion Claude and Gemini demonstrate moderate to strong agreement with experts for cognitive level classification and detection of objective technical flaws, suggesting their potential as adjunctive tools in MCQ review. However, weak agreement on learning outcome alignment and variability across models indicates that LLMs cannot yet replace expert judgment. A hybrid approach combining LLM-assisted screening with human expertise may optimize item quality assurance in medical education. These findings derive from a single institution and discipline with a limited item set (n = 85) and require multi-center validation before broader generalization.

M. O. Aydin, B. Coşkun, İbrahim Hamal · 0 citations
Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting. In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1–21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used. Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran’s Q = 14.089, df = 4, p  = 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted p  = 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all p  < 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses. Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.

Yan-Ru Jiang, Qianyun Wang, Liang Zheng et al. · 0 citations