Skip to content
Open access

LLM-as-a-judge for infection prevention and control and antimicrobial resistance impact: comparing three main LLMs vs. human experts' assessment

Jul 2026 · Frontiers in Public Health · Vol 14 · 2 citations · 31 references
Medicine

TL;DR

LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators, and are best suited as a scalable screening tool within supervised human-in-the-loop workflows.

Abstract

Background Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks. Methods We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines. Results Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions. Conclusions LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.

Read PDF

Similar papers

Review Open access Sep 2026

Large language models for clinical decision support in infectious disease diagnosis and antimicrobial prescribing: a scoping review

This scoping review is the first scoping review to focus specifically on large language models at the intersection of infectious-disease diagnosis and antimicrobial prescribing, rather than on artificial intelligence in medicine broadly.

M. Sannathimmappa · 0 citations
Review Open access Sep 2026

Defining use cases for biomarkers and tests across tuberculosis infection, disease and treatment: An international consensus and prioritisation exercise

Background Translation of tuberculosis (TB) biomarker and diagnostic research into tools that improve patient and public health outcomes has been slow, partly because no internationally agreed framework exists defining the use cases that new biomarkers and tests should address. We aimed to identify, validate, and prior...

S. Wright, F. Fama, A. de Wilton et al. · 0 citations
Aug 2026

Testing Knowledge Boundaries: Adversarial Evaluation of LLMs for Antimicrobial Stewardship.

Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk, and endorsed LLMs as useful AMS support tools with moderate supervision.

Ángela Abejez-Arrizabalaga, Galadriel Pellejero-Sagastizabal, Rocío Aznar-Gimeno et al. · 0 citations
Open access Aug 2026

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are eva...

Qi-Qi Zheng, Ru Chen, Ming-Ming Cai et al. · 0 citations
Review Open access Sep 2026

A Scenario-Based Survey of Clinician Recommendations for Follow-Up Blood Cultures in Patients With Varying Risk of Persistent Gram-Negative Bacteremia

Abstract Background The utility of follow-up blood cultures (FUBCs) in patients with gram-negative bacteremia (GNB) is unsettled. In practice, the use of FUBCs is variable. Understanding practice variation and drivers of FUBC ordering may clarify how clinicians identify patients at high risk for persistent GNB. Methods...

Roberta Monardo, Halie L. Hotchkiss, Rebecca North et al. · 0 citations
Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.