Review
Aug 2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, it is found that a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges are found.
M. Flathers, Phuong Anh Nguyen, J. Noorily et al.
· 0 citations