Skip to content
Review Open access

Large Language Model Benchmarks in Medical Tasks

Oct 2024 · Medicine Advances · Vol 4, pp. 316 - 341 · 32 citations · ⚡ 1 influential · 366 references
Computer Science

Abstract

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care III (MIMIC‐III), MIMIC‐IV, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.