Skip to content
Open access

Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment

Sep 2026 · PLOS Digital Health · Vol 5 · 0 citations · 40 references
Medicine

Abstract

Radiology reports are often filled with medical jargon that limits patient understanding. Lay summaries can improve understanding but are time-consuming for healthcare providers to create. The objective of this study is to explore the use of tailored prompts for five Large Language Models (LLMs) in generating lay summaries from radiology reports. Using 100 reports from the publicly available “BioNLP 2023 report summarization” dataset, lay summaries were generated by each LLM, under select prompting styles [Few-Shot (GPT-4), Generated Knowledge (GPT-4o mini, Gemini 1.5 – Pro, Gemini 1.5 – Flash), and Zero-Shot (Llama 3.1)] informed by a pilot work. The summaries were evaluated using a mixed-method framework: subjective assessment (Likert statements) by blinded experts (n = 2 radiology fellows) and Large Reasoning Models (LRMs) [(Gemini 2.5 – Pro (LRM 1); GPT-oss-120b (LRM 2)], and readability metrics (Flesch-Kincaid Grade Level and Flesch Reading Ease). Using percentage agreement of Likert statements, the LLM-prompt combinations’ performances were ranked, and Friedman and post-hoc Nemenyi tests were conducted. Gemini 1.5 - Flash and - Pro (generated knowledge) were rated highest by human experts and LRMs for generating actionable lay summaries that require minimal supervision [P < 4.97 × 10-2 (Rater 1); P < 9.03 × 10-21 (Rater 2), P < 6.90 × 10-15 (LRM 1), P < 2.760 × 10-5 (LRM 2). GPT-4 (few-shot) achieved the highest human-rated accuracy (98%), while Gemini 1.5 – Flash (LRM 1-rated: 95%) and Gemini 1.5 – Pro (LRM 2-rated: 91%) ranked first in LRM-rated accuracy. Gemini 1.5 - Pro produced the most accessible summaries (Flesch-Kincaid Grade Level: 7.55 ± 1.38, Flesch Reading Ease: 67.84 ± 7.78). Strong agreement was observed between experts and LRMs [0.96% (LRM 1) and 3.4% (LRM 2) complete disagreement]. Overall, this study highlights Gemini-models with generated knowledge prompts and the potential of LRM evaluators in assessing LLM-generated lay summaries.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.