Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting....