Qualitative analysis revealed that empathy, exploratory questioning, contextual understanding, and personalization were key factors shaping users’ trust, suggesting design implications for LLM-based counseling systems.
Abstract
Large language models (LLMs) have expanded the potential of conversational AI in mental health support, yet counseling inherently relies on trust and relational aspects that may not transfer directly to these systems. We examine how users’ trust and experiences differ between human and LLM-based counseling, conducting a within-subjects study in which participants discussed their own concerns with both a licensed human counselor and an LLM-based counselor through text-based sessions. We assessed subjective distress and trust across five dimensions, and analyzed open-ended feedback across different trust profiles. The human counselor condition received higher overall trust, with the largest gaps in Faith and Personal Attachment, while Understandability remained comparable across conditions. Participants also reported greater reductions in subjective distress following human counselor sessions. Qualitative analysis further revealed that empathy, exploratory questioning, contextual understanding, and personalization were key factors shaping users’ trust, suggesting design implications for LLM-based counseling systems.
The rapid expansion of large language models (LLMs) has created new opportunities for university mental health services, particularly in contexts where counseling demand exceeds available professional resources. This study examined the application of an LLM-based intelligent agent in Chinese university counseling settings, focusing on three key outcomes: perceived counseling alliance, disclosure willingness, and risk recognition. Using a randomized between-subjects experimental design, 388 valid responses were collected from university students and assigned to either an LLM-based counseling condition or a control condition. The LLM-based agent significantly improved perceived counseling alliance and disclosure willingness, and perceived counseling alliance partially mediated the relationship between condition and disclosure. Scenario-based analyses further showed that the LLM-based agent improved general risk recognition and slightly enhanced sensitivity to high-risk cues, although performance remained more limited for crisis-level disclosures. These findings suggest that LLM-based agents are most effective as front-end support tools that facilitate engagement and emotional expression rather than as autonomous crisis detectors. In the context of Chinese university counseling, the study supports a complementary human–AI model in which intelligent agents lower barriers to help-seeking while trained counselors retain responsibility for risk assessment and intervention. Overall, the results contribute to digital mental health research by clarifying both the relational benefits and safety boundaries of LLM-supported counseling.
Trust in AI for emotional support is not universal; it is shaped by who users are, where they come from, and what they value. Yet research in this area lacks validated psychometric instruments for assessing user perceptions in affective AI contexts and large-scale evidence on how trust formation varies across user segments. To address these gaps, we develop and validate a seven-construct psychometric scale, test a Structural Equation Model (SEM) linking system attributes to Trust and Perceived Benefits as mediators of Actual System Use, and conduct a Multi-Group Analysis (MGA) across five sociodemographic dimensions (gender, age, education, socioeconomic status, cross-national region), drawing on 1,343 active users from seven countries. We find that users experience empathy and anthropomorphism as a unified"Humanlikeness"construct, and that Privacy, Personalization, and Humanlikeness drive Trust while Perceived Bias degrades it. Notably, adoption logic diverges across groups: Privacy shapes women's trust more than men's, Anglosphere (UK, USA) users respond more positively to Humanlikeness than Europeans, and educated and higher-income users require Trust to engage, whereas older adults and lower socioeconomic groups bypass it entirely, relying on perceived practical benefits (e.g., 24/7 availability, non-judgmental support). Our findings extend technology acceptance theory and inform the equitable design of emotional support AI.
Natalia Amat-Lefort, M. Yazan, A. C. Curry et al.· 0 citations
Research on physician communication and patient trust has expanded in recent years, yet the evidence remains fragmented across conceptualizations of communication, operationalizations of trust, research methods, and clinical and technological contexts. This systematic review synthesizes recent studies examining how physician communication styles are associated with patient trust in real physician-patient interactions. Following PRISMA guidelines, 13 studies published between January 2021 and July 2025 were identified from 5,730 records retrieved from English-language peer-reviewed journals. Across quantitative, qualitative, mixed-methods, and review-based evidence, physicians' communication styles-including information-giving, socio-emotional, partnership-building, paternalistic, and nonverbal approaches-were examined in relation to patient trust. Overall, clear information delivery, empathy, active listening, shared decision-making, and supportive nonverbal behaviors were generally associated with higher levels of patient trust, while directive or paternalistic communication showed more context-dependent effects. Using established multidimensional understandings of patient trust as an analytic lens, the review shows that communication-trust associations were shaped by clinical settings and digital or technological conditions. Taken together, the findings highlight key patterns and gaps in the literature and underscore the need for more rigorous and cross-cultural research, as well as communication training that emphasizes empathy and partnership to strengthen physician-patient trust.
Zhaoyang Tan, Zhengnan Sun, Yujun Lin· Health Communication· 0 citations
The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.
Shahrul Hazman Shamshudeen, N. Sharef, Muhamad Saiful Bahri Yusoff· Journal of Information &...· 0 citations
While ChatGPT was primarily viewed as an efficient tool in a work context, the quantitative survey reveals a weakly significant correlation between psychological stress and openness toward the social-emotional use of chatbots.
Stefanie Osetrow, H. Klapperich, Dr. Alina Huldtgren· Proceedings of Mensch und Co...· 0 citations
Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.
Alexander Marrapese, Basem Suleiman, Jinglin Sun et al.· International Conference on...· 0 citations