Skip to content

Author

Joshua Davis

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Development of a benchmarking dataset for symptom detection using large language models

Abstract Objectives To develop a pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters. Materials and Methods We created a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts (264 encounters; 16 symptoms; double-coded and adjudicated). Nine different LLMs from 4 vendors (OpenAI, Meta, DeepSeek, Moonshot AI) were used as examples to test our evaluation pipeline; outputs were assessed for correct structure and symptom information. Results Of 3085 excerpts, 2087 (68%) contained symptoms. Pain, cough, and shortness of breath were most common; LLMs achieved F1 scores ranging 0.66-0.88 for these symptoms with minimal prompt engineering. Of tested models, GPT-4.1 demonstrated the best overall performance. Discussion Our evaluation pipeline and benchmarking dataset are publicly available and applicable to various LLMs, including open-source models. Conclusion This work supports the development and optimization of models that seek to improve patient symptom understanding.

Joshua Davis, B. Durieux, C. V. van Dongen et al. · 0 citations