Skip to content
Review Open access

Large Language Model versus Clinician Written Summaries of Research Papers.

Sep 2026 · Journal of the American Board of Family Medicine · Vol 39 1 · 0 citations
Medicine

TL;DR

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.

Abstract

INTRODUCTION Clinicians require concise, accurate summaries of new research to inform practice. Patient-Oriented Evidence that Matters (POEMs), published in American Family Physician, are a benchmark for summarizing primary literature in family medicine, while large language models (LLMs) offer scalable summarization but require rigorous evaluation. The objective of this study was to evaluate the accuracy and quality of summaries generated by large language models compared with expert-authored POEMs.

Methods

In this study, we compared LLM-generated summaries (Microsoft Copilot, GPT-4o class) with 24 recent matched POEMs using a standardized prompt. Two trained raters independently scored each summary with a 13-item tool (score range 0-13), cataloged errors, recorded word counts, and indicated preferences on a 5-point scale.

Results

LLM summaries outperformed POEMs in total score (mean 12.1 vs 10.6; mean difference 1.5, 95% CI 1.1-2.0; P < 0.001), with similar lengths (328 vs 353 words; P = 0.23). Errors occurred in fewer LLM-DOCSs (2/24) than POEMs (9/24), with a mean error score difference of 20% (95% CI 7% -33%; P < 0.001). POEMs most often missed in the categories Contextual Background and Limitations; both approaches frequently missed in Clinical Applicability. Reviewer preference favored LLM-DOCS (mean 2.44 on a 1-5 scale; 95% CI 2.1-2.8).

Conclusions

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact. Findings support pragmatic LLM-assisted summarization and highlight the need for standardized evaluation tools and explicit prompts for clinical applicability.

Read PDF

Similar papers

Open access Sep 2026

Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment

Radiology reports are often filled with medical jargon that limits patient understanding. Lay summaries can improve understanding but are time-consuming for healthcare providers to create. The objective of this study is to explore the use of tailored prompts for five Large Language Models (LLMs) in generating lay summa...

Nanziba Tasneem, C. B. van der Pol, A. Zahoor et al. · 0 citations
Review Open access Aug 2026

Trial Files: Leveraging large language models to summarize practice-changing clinical trials for clinicians

Background Each day, over 100 randomized controlled trials (RCTs) are published, making it impossible for clinicians to stay up-to-date with medical literature. Large language models (LLMs) can identify and summarize emerging clinical evidence and support medical education. Methods We created and prospectively evaluate...

Katarina Zorcic, E. Bartsch, Bryant Lim et al. · 0 citations
Review Open access Sep 2026

Citation reliability of frontier large language models in medical writing and its automated verification

Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash, generating 270 cardiology narrative reviews with web search enabled, a...

R. Shin, J.-M. Lee, J. Park et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations,...

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
Open access Sep 2026

Comparison of physician-authored and artificial intelligence-generated after-visit summaries: A blinded comparative study.

In this blinded evaluation, LLM-generated AVSs were clearer, more comprehensive, and more empathetic than physician-authored AVSs and were associated with lower physician-rated potential for harm.

Milla Kviatkovsky, Caden Stewart, Annalise J. McDonald et al. · 0 citations
Review Open access Sep 2026

Large language models for patient-facing pathology report interpretation: A scoping review.

OBJECTIVE Pathology reports are increasingly released directly to patients, but their diagnostic language is primarily designed for clinicians. This scoping review examined how large language models (LLMs) have been used for patient-facing pathology report interpretation, how generated outputs have been evaluated, and...

Chen Wang, Jie Hao, Si-Jia Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.