Skip to content
Open access

Benchmarking 17 Large Language Models Against Dermatology Residents: A Comparative Study on Board-Style Medical Knowledge

Sep 2026 · Dermatology Practical & Conceptual · 0 citations · 8 references

Abstract

Introduction: Large Language Models (LLMs) show promise in medical domains, yet benchmarks comparing diverse LLM architectures with dermatology residents remain scarce. Objectives: To compare the performance of 17 LLMs against dermatology residents on a board-style examination and to evaluate performance differences across model families and residency seniority levels. Methods: An exploratory cross-sectional comparative study employed a 50-item, single-best-answer, text-based examination covering core dermatological topics. Each correct answer was awarded 2 points (maximum score: 100). Seventeen LLMs (OpenAI, Google, Anthropic, and open-source/other models) and 14 dermatology residents (postgraduate years 1–4) were assessed. Statistical analyses included Mann–Whitney U, Kruskal–Wallis, Spearman correlation, and Fisher’s exact tests, with sensitivity analyses addressing within-family correlation among LLMs. Results: LLMs significantly outperformed residents (median: 94.0 [IQR: 90–96] vs. 76.0 [IQR: 74.5–82]; U = 225.0, p < 0.001; r = 0.89; Hodges–Lehmann median difference: 16.0 points, 95% CI: 12.0–20.0). Sensitivity analyses using one representative model per family confirmed this finding (p = 0.003). While 82.4% of LLMs scored ≥90, no residents reached this threshold (p < 0.001). A statistically significant difference was observed among model families (Kruskal–Wallis H = 8.86, p = 0.031); however, post-hoc pairwise comparisons did not reach significance after Bonferroni correction. No significant correlation was found between residency seniority and exam scores (ρ = 0.29, p = 0.31). Conclusions: In this exploratory single-center study, current-generation LLMs demonstrated near-optimal performance on a text-based theoretical dermatology examination, significantly exceeding resident scores irrespective of seniority. These models may serve as supplementary educational tools for board preparation; however, clinical applicability requires further evaluation through multimodal, image-based assessments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.