Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot
Abstract
Background Patients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues. Methods This cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28–29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons. Results Interrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p < 0.001), and for GQS (p = 0.011). Within this defined late-May 2026 response set, responses returned by the interface displaying the label ‘Grok-4.3’ had the highest observed sample means for DISCERN (58.68), EQIP (81.14), GQS (4.27), and the JAMA visible metadata/transparency proxy score (1.00). No evaluated public-interface response set achieved recommended sixth-grade readability, and no individual response met all six readability thresholds. Responses returned by the interface displaying the label ‘DeepSeek-v4’ had the most favorable observed readability profile within this sampled response set, although the mean FKGL remained 11.53. These observations should not be interpreted as durable model rankings. Conclusion In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag. Pronounced ceiling and floor effects preclude conclusions about clinical sufficiency or safety. These findings should not be generalized to non-English use, different health-literacy levels, country-specific emergency-care pathways, other regions or account configurations, or later interface states.