Jul 2026· International Conference on Digital Health· pp. 237-246· 0 citations· 24 references
Abstract
Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.
Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program"Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman et al.· 0 citations
Mental health (MH) chatbots are increasingly used to provide accessible, on-demand emotional support, yet it remains unclear how these systems linguistically construct and communicate care. This work-in-progress examines whether MH chatbots produce responses that reflect supportive value orientations and counseling-adjacent tone. We conduct an observational analysis of responses from three widely used MH chatbots (Wysa, Sintelly, and Youper) across context-aware scenario prompts and a standardized-question session. Responses are analyzed using the SemEval’23 “Adam Smith” human value detection model and LIWC’22 psycholinguistic measures, including Language Style Matching (LSM), Clout, and Authenticity. Values such as “Security: Personal” and “Benevolence: Caring” appear consistently across systems, with contextual variation in secondary value emphasis. Linguistic patterns show moderate-to-high LSM and consistently high Clout, with Authenticity varying by scenario. These findings are exploratory signals intended to inform future evaluation and design of supportive conversational mental health systems.
Maleeha Sheikh, Chao Chen, Romael Haque· Information Hiding· 0 citations
The advent of Large Language Models has accelerated interest in empathetic conversational agents. Despite a surge in empirical research, artificial empathy remains deeply fragmented, often serving as a catch-all term for diverse interactional phenomena. Addressing this conceptual gap, we systematically review 89 empirical studies to map how human-machine empathy is operationalized. Our synthesis reveals that empathy is highly situated and driven by functional goals, like health and well-being, transactional service, social interaction, and learning support. Within these contexts, we classify affective responsiveness by its directional flow, detailing how agents project, elicit, or mediate empathy. We structure the literature into a cohesive framework spanning linguistic, paralinguistic, identity, and architectural strategies. Furthermore, our methodological evaluation reveals a reliance on adapted clinical metrics, a scarcity of longitudinal studies, and a disproportionate focus on text-based over voice-based interfaces. Ultimately, this review equips researchers and practitioners with an actionable foundation for designing, measuring, and implementing contextually appropriate and empathetic agents.
Supriya Khadka, Smit Desai· International Conference on...· 0 citations
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
P. A. Fonseca, R. Rodríguez-Carvajal, Rafael A. Calvo· 0 citations
Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language models (LLMs), improving cooperation and reducing ethical violations in LLM outputs. However, as new models, evaluation metrics, and benchmarks emerge rapidly, it remains challenging to systematically assess whether and how contemplative principles enhance LLM alignment across diverse and evolving scenarios, and existing approaches are often ad hoc and fail to generalize. We present a modular, extensible evaluation framework, initially targeted at the mental health domain, that enables seamless integration of new models, metrics, and benchmarks through a reusable pipeline. The framework currently reproduces existing state-of-the-art results and supports systematic cross-evaluation by flexibly mixing and matching models, metrics, and benchmarks, enabling fair comparison and deeper insight. Its plug-and-play prompting module offers a principled pathway for incorporating ethical perspectives such as contemplative principles, allowing domain experts to define alignment criteria without requiring technical expertise. Although initially focused on mental health, the framework is domain-agnostic and extends naturally to areas such as decision-making, moral reasoning, and human-AI collaboration. By bridging computational evaluation with human-centered ethical reasoning, this work lays the groundwork for interdisciplinary research spanning cognitive science, behavioral economics, philosophy, and system design, toward robust, trustworthy, and socially beneficial human-AI ecosystems.
Asher Sprigler, Yang-Yang Feng, Iftach Amir et al.· 0 citations
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos et al.· 0 citations