Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation
Abstract
To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions. In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers. The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7–4.0] vs. 2.7 [2.0–3.7], P < 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7–4.3] vs. 2.7 [2.3–2.7], Δ median = −1.3, P < 0.001; conventional models: 2.7 [2.0–3.7] vs. 2.3 [2.0–2.7], Δ median = −0.4, P < 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type. Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.