Skip to content
Review

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, it is found that a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges are found.

Abstract

General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($\tau \ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

View source

Similar papers

Open access Jul 2026

PsyEval: a comprehensive large language model evaluation benchmark for mental health.

This work introduces PsyEval, a benchmark specifically designed to evaluate LLMs in mental health-related tasks across three core dimensions: knowledge, diagnosis, and emotional support, and reveals considerable gaps in LLMs' current ability to reason accurately and respond appropriately in mental health contexts.

Haoan Jin, Siyuan Chen, Dilawaier Dilixiati et al. · 0 citations
Conference Jul 2026

Trustworthy Mental Health Assessment via Confidence-Guided LLMs

Depression and anxiety disorders are among the most prevalent and debilitating mental health conditions worldwide, imposing substantial personal, social, and economic burdens. Although recent advances in Large Language Models (LLMs) have shown promise in supporting mental health assessment and intervention, existing approaches often lack contextual awareness, real-time adaptability, and privacy-preserving personalization. To address these limitations, we propose a novel, context-aware and privacy-preserving mental health evaluation architecture that synergistically integrates LLM-driven intelligence. The proposed system enables personalized, continuous, and stigma-free mental health support by combining structured multiple-choice questionnaires with advanced language models, including GPT-3.5-turbo and Groq, to analyze user inputs, identify behavioral patterns, and predict potential mental health conditions such as depression and anxiety. Furthermore, the platform provides individualized recommendations, including self-care strategies, lifestyle adjustments, mindfulness practices, and referrals to healthcare professionals when appropriate. Recognizing the critical importance of reliability in sensitive healthcare settings, we introduce an ensemble-based aggregation framework that explicitly incorporates classification confidence and uncertainty quantification across multiple LLMs. Experimental results demonstrate that the proposed approach outperforms existing LLM models. By prioritizing user anonymity and data privacy, the proposed system reduces psychological barriers to seeking mental health support and promotes early intervention.

Jashraj Jani, Sara Akif, Wassila Lalouani · 0 citations
Review Aug 2026

Large language model applications for real-time clinical mental health assessment: Current potential and future directions.

Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.

K. Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al. · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Open access Jul 2026

A bibliometric analysis of large language models in mental health research

The analysis revealed a rapid acceleration in scholarly output, with a compound annual growth rate of 140%, driven by advancements in models such as GPT-3 and GPT-4, alongside strategic funding and industry initiatives.

Mohammad Ali Hussiny, T. Saidi, Minna Pikkarainen et al. · 1 citation