Skip to content
Review Open access

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Aug 2026 · JMIR AI · Vol 5 · 0 citations · 123 references
Medicine

Abstract

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Read PDF

Similar papers

Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations
Review Open access Jul 2026

Natural Language Processing Applied to Psychiatric Clinical Notes: Scoping Review

Abstract Background Psychiatric clinical notes in electronic health records (EHRs) provide rich longitudinal information that can support clinical decision-making. Using historical medical data can enable earlier identification of mental illness, better characterization of disease trajectories, and more personalized treatment planning. Natural language processing (NLP) transforms these unstructured notes into analyzable representations for research and care. Objective This study aims to systematically summarize NLP methodologies for psychiatric clinical notes, compare major modeling paradigms and application areas, and highlight emerging large language model (LLM) trends, key challenges, and future research directions. Methods Following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines, a literature search was conducted for articles on NLP methods based on psychiatric clinical notes published from January 2021 to December 2025 in Ovid MEDLINE, Ovid EMBASE, PubMed, Scopus, Web of Science, the ACM Digital Library, and ScienceDirect. This scoping review analyzed NLP methods applied to psychiatric clinical notes, focusing on major trends, identifying suitable features for traditional machine learning (ML)–based models, applications of pretrained language models (PLMs), and key challenges. Approaches were categorized as rule-based, traditional ML, hybrid, deep learning (DL), and LLM-based methods across information extraction and text classification tasks. Results In total, 101 studies were eligible for inclusion. Rule-based methods (n=36) and hybrid approaches (n=34) remained the most widely used techniques, largely favored for their interpretability in handling nuanced, subjective clinical notes. These were followed by DL (n=15), traditional ML (n=10), and LLM-based approaches (n=6). Traditional ML studies relied heavily on engineered features, which could be grouped into 5 broad categories: domain knowledge features, lexical and statistical features, vector-based semantic features, emotion-related features, and temporal features. PLMs improved performance mainly through domain adaptation and task-specific fine-tuning, enhancing the handling of psychiatric language, medical terminology, and clinical note structure. LLM-based studies, although still limited in number, indicated a growing shift toward generative and reasoning-based applications. Conclusions Hybrid NLP approaches remain dominant, combining domain rules with ML for extraction and classification. DL approaches continue to advance, with domain adaptation supporting medical terminology and clinical semantics. LLMs may further automate complex workflows via zero-shot capabilities and reasoning, alongside growing interest in temporal modeling and multimodal integration. Key future needs include improved generalizability across institutions, privacy protection, and careful attention to ethical implications in clinical deployment.

Shuying Rao, Xiangye Chen, Guifeng Deng et al. · 0 citations
Review Open access Aug 2026

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al. · 0 citations
Open access Jul 2026

A bibliometric analysis of large language models in mental health research

The analysis revealed a rapid acceleration in scholarly output, with a compound annual growth rate of 140%, driven by advancements in models such as GPT-3 and GPT-4, alongside strategic funding and industry initiatives.

Mohammad Ali Hussiny, T. Saidi, Minna Pikkarainen et al. · 1 citation