Background Artificial intelligence (AI) is increasingly used for mental health support, through both purpose-built products and general-purpose large language models (LLMs), at a pace far greater than evidence of their efficacy or safety. There is no consensus taxonomy of the potential harms arising from this use, and existing classification frameworks are fragmented across clinical, computational, ethical, and regulatory domains. Objective This scoping review aims to systematically map the harms which may arise from engaging with AI for mental health support, as formulated in the academic literature, regulatory frameworks, grey literature, and practitioner guidance, to inform the co-production of a harm taxonomy. Methods We will follow the framework of Arksey and O’Malley, as refined by Levac et al. and the Joanna Briggs Institute, with reporting adhering to the PRISMA Extension for Scoping Reviews (PRISMA-ScR). We will search MEDLINE (Ovid), Embase (Ovid), APA PsycINFO (EBSCOhost), Scopus, Web of Science, ACM Digital Library, IEEE Xplore, ProQuest Dissertations and Theses Global, PROSPERO, and Google Scholar for peer-reviewed literature and preprints, combining terms for AI technologies, mental health, and harm. Grey literature searches will cover regulatory bodies, professional associations, the Mental Health Innovation Network, OpenSyllabus, and preprint repositories. No date restriction will be applied; searches will be restricted to English. Titles, abstracts, and full texts will be screened independently by two reviewers following a calibration exercise. Data will be charted using a piloted standardised instrument and synthesised narratively through descriptive tabulation and iterative harm mapping, with consultation of lived experience advisors and clinical experts before finalisation. Ethics and dissemination No primary data will be collected, so ethical approval is not required. Findings will be published as a peer-reviewed manuscript, presented at conferences, and used to inform co-production workshops within the SAFER-MH programme.
X. Hunt, A. G. Mokaya, Sara Zannone et al.· Wellcome Open Research· 0 citations
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers'safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
Evaluating generative AI output remains a critical bottleneck for safe and scalable deployment of AI in healthcare. Expert clinical judgement is often presented as the gold standard, but human assessment is costly and inconsistent. LLM-as-judge systems, i.e., leveraging AI to evaluate other AI outputs, have been proposed, yet their reliability in global health remains untested. We compared five LLM judges and six human clinicians in evaluating responses to questions posed by Rwandan health workers. The highest-performing LLM-judge (Claude-4.1-Opus) matched human evaluators on only four of eleven evaluation criteria, with other models scoring too leniently (Gemini-2.5-Pro) or too harshly (GPT-5). Constructing LLM-juries to balance model-specific biases improved agreement on only one additional criterion. Notably, performance and cost-effectiveness fell when moving from English to Kinyarwanda. Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.
G. Williams, S. Rutunda, Floris Nzabakira et al.· npj Digital Medicine· 1 citation
A novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria reveals several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts.
Tobi Olatunji, C. Aka, C. Okocha et al.· medRxiv· 0 citations