Assessing multiple-choice question quality in internal medicine: a comparative analysis of three large language models against expert consensus
Background Large language models (LLMs) are increasingly explored for their potential to support quality assurance in medical education assessment. However, limited evidence exists on the alignment between LLM evaluations and expert judgment across multiple dimensions of multiple-choice question (MCQ) quality. Methods This comparative methodological study evaluated 85 MCQs from an internal medicine clerkship examination. Three LLMs (Claude Sonnet 4, Gemini 2.5 Flash, and Llama 3.3 70B Instruct Turbo) and three medical education experts independently assessed each question for cognitive level (Revised Bloom’s Taxonomy), alignment with the intended learning outcome (5-point Likert scale), and presence of technical flaws based on NBME guidelines. Agreement was calculated using Fleiss’ kappa for cognitive level classification and Cohen’s kappa for binary technical flaw criteria, with intraclass correlation coefficients (ICC) for Likert-scale alignment ratings. Results For cognitive level classification, Gemini (κ = 0.424, p < 0.001) and Claude (κ = 0.415, p < 0.001) showed moderate agreement with experts; Llama demonstrated lower agreement (κ = 0.266, p < 0.001). Alignment with learning outcomes yielded weak-to-moderate agreement for all models (ICC 0.121–0.380). For technical adequacy, Claude and Gemini achieved almost perfect agreement on detecting negatively worded stems (κ = 0.897, p < 0.001) and substantial agreement on inconsistent numerical data (κ = 0.661, p < 0.001) but showed poor agreement on more subjective flaws. Llama performed poorly across most technical criteria. Conclusion Claude and Gemini demonstrate moderate to strong agreement with experts for cognitive level classification and detection of objective technical flaws, suggesting their potential as adjunctive tools in MCQ review. However, weak agreement on learning outcome alignment and variability across models indicates that LLMs cannot yet replace expert judgment. A hybrid approach combining LLM-assisted screening with human expertise may optimize item quality assurance in medical education. These findings derive from a single institution and discipline with a limited item set (n = 85) and require multi-center validation before broader generalization.