Large Language Model Performance on Multistep Clinical Cases: Comparative Study Across Question and Case Levels.
BACKGROUND Most large language models (LLMs) have achieved passing scores on medical licensing examinations. However, most evaluations focus on single-question accuracy, overlooking performance on multistep patient management scenarios, such as making a diagnosis followed by a treatment plan. It is unclear if LLMs can...