A model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions is developed, providing a scalable framework for safer deployment across models.
Abstract
Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers'safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
Anian is presented, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulness-intervention routing that supports the internal feasibility of the label framework and gating logic but does not establish clinical validity, diagnostic accuracy, real-world safety, or effectiveness.
Treating “safety” as synonymous with toxicity prevention made sense when AI systems were passive query responders. It makes far less sense once those systems handle mental health triage or customer grievance escalation, where reading and matching emotional register is a core functional requirement. RLHF-based alignment penalises expressive, affectladen outputs and drives models toward a neutral mean. We hypothesise that sustained alignment pressure progressively compresses a model's emotional output space, a phenomenon we term Emotional Flattening. To examine this, we use Activation Steering as a controlled proxy and track degradation in Gemma2B and Llama-3-8B across increasing steering strengths. The failure modes are architecture-dependent. Gemma-2B collapses into degenerate repetitive output; Llama-3-8B retains syntactic coherence while losing affective signal. Both outcomes erode the relational competence proactive agents depend on, pointing to what we term the Social Safety Gap in current alignment evaluation.
Ajitesh Sharma, Rayban Pranav Mahesh, Aarya Ashish Nagvekar et al.· Annual International Compute...· 0 citations
Large language models (LLMs) are rapidly becoming embedded in everyday mental health help-seeking practices, particularly among young people who already turn to digital platforms as gateways to mental health support. While LLMs offer unprecedented immediacy and accessibility, their integration into help-seeking ecosystems raises important questions for digital health research. This opinion paper argues that LLMs fundamentally reshape the developmental processes underpinning online help-seeking. Traditional digital help-seeking requires active exploration, searching, comparing sources and reflecting on lived experience, processes that contribute to mental health literacy and resilience. In contrast, LLMs collapse informational plurality into singular, authoritative-sounding responses, potentially shifting users from active exploration toward passive consumption. We discuss the risks of sycophancy, and over-reliance on immediacy, and consider how these dynamics may alter developmental trajectories of coping and help-seeking agency. We argue that preserving agency, connectedness, and reflective engagement must be central to the design of conversational AI in health contexts.
Claudette Pretorius· Information Hiding· 0 citations
The integration of generative artificial intelligence (AI) and large language models (LLMs) into healthcare has accelerated dramatically, with emergency medicine emerging as a particularly dynamic yet challenging domain for clinical deployment. This comprehensive narrative review synthesizes contemporary literature to examine current clinical applications, critical safety considerations, implementation scalability and governance, and future research priorities surrounding generative AI in acute care settings. We examine the current clinical applications, critical safety considerations, implementation scalability and governance, and future research priorities for generative AI in acute care settings. Current emergency medicine applications span automated documentation via ambient clinical intelligence systems, clinical decision support for triage and diagnostic prediction, patient communication tools including discharge summary generation, and multilingual data extraction supporting care transitions. Despite promising efficiency gains - including documented reductions in clinician burnout from 50.6% to 29.4% and modeled documentation time savings of up to 7.1 hours per shift cycle - substantial safety concerns persist, with hallucination rates ranging from 26% to 36% across automated pipelines and systematic misclassification in high-acuity triage tasks. We address automation bias (26% increased risk) and data privacy and governance risks, and identify algorithmic equity as a critical research priority. We conclude that while generative AI holds transformative potential to reduce clerical burden and augment clinical reasoning, its successful deployment in emergency medicine requires rigorous attention to clinical safety, health equity, workflow integration, and human‑factors considerations.
Klaudia Kwolek, Wiktoria Laskowska, M. Pilarek et al.· International Journal of Inn...· 0 citations
Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program"Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman et al.· 0 citations