Skip to content
Review Open access

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

Sep 2026 · Radiology Advances · 0 citations

TL;DR

This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).

Abstract

High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create. To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses. Fifty Radiopaedia cases across 10 subspecialties were used to generate 200 radiology OSCE-style sessions (three questions per session) using four LLMs: GPT-4o (4O), Llama 3-70b (LM), Claude 3.5 Sonnet (CL), and Gemini 1.5 Flash (GM). Three expert radiologists blindly evaluated sessions for clarity, clinical relevance, difficulty, option accuracy, assessment accuracy, and feedback quality on 5-point Likert scale. Generalized estimating equations (GEE) compared model performance, accounting for rater correlation. The primary accuracy metric was Top Two Box Accuracy (2TBA; all raters ≥4); secondary metrics included Top Box Accuracy (TBA), Average Score ≥4 Accuracy (AS4A), and Perfect 5 Accuracy (P5A; all raters = 5). Final rankings were determined via Borda count. Inter-rater reliability ranged from fair to good agreement (Gwet's AC2: 0.39–0.73). GEE analysis showed significant performance variations (p < 0.05) across metrics. Final Borda scores: 4O (18.5), LM (13.5), CL (11.0), GM (7.0). TBA was high (70.7–100%). For the primary metric 2TBA, 4O showed significantly higher odds of achieving high-quality scores vs. GM for clarity (odds ratio [OR] 4.52, p < 0.001) and overall assessment accuracy (OR 3.78, p < 0.001). Under the strictest threshold (P5A), 4O led in composite Question General Accuracy (52.7% vs. 32.7% LM, 21.3% CL, 1.3% GM). Qualitative failure modes included distractor ambiguity, conceptual repetition, and information leakage. This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024). 4O demonstrated the highest relative performance, though none consistently achieved expert-level quality. While LLMs can augment radiology education, expert review remains mandatory for high-stakes assessment.

Read PDF

Similar papers

Open access Sep 2026

Do Large Language Models Use the Clinical Vignette? A Question Ablation Study on the Orthopaedic In-Training Examination

Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...

F. Gafoor, M. Syed, M. Halai et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Sep 2026

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.

K. Ulutaş, A. Pekmezci · 0 citations
Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

Gemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education, however, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy.

Burak Altunpak · 0 citations
Review Aug 2026

Generating Image-Based Multiple-Choice Questions with Multimodal Large Language Models: Expert and Psychometric Evaluation.

Multimodal large language models are best positioned as supervised tools requiring radiologist verification and item-level psychometric evaluation in multiple-choice questions generated by GPT-4o and o3 from correctly recognized radiographs and compare selected items with faculty-written questions.

E. Emekli, Murat Tepe, Serhat Demir et al. · 1 citation
Open access Sep 2026

Benchmarking 17 Large Language Models Against Dermatology Residents: A Comparative Study on Board-Style Medical Knowledge

Introduction: Large Language Models (LLMs) show promise in medical domains, yet benchmarks comparing diverse LLM architectures with dermatology residents remain scarce. Objectives: To compare the performance of 17 LLMs against dermatology residents on a board-style examination and to evaluate performance differences ac...

Mahir Dığış, Kısmet Kaya, B. Demir · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.