Skip to content

Performance of a domain-specific large language model in answering patient questions in psychiatry

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Abstract

Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.

View source

Similar papers

Open access 2026

MOST FREQUENTLY ASKED QUESTIONS BY OLDER ADULTS IN GERIATRIC REHABILITATION: EVALUATING LARGE LANGUAGE MODELS AS A SOURCE OF INFORMATION

Introduction: Older adults in geriatric rehabilitation are increasingly turning to internet-based resources and large language models for healthrelated information outside of clinical follow-up. The aim of this study was to compare the reliability, clinical accuracy, quality, usefulness, and readability of ChatGPT-5.2, Gemini 3, and DeepSeek V3.2 responses to patient questions regarding geriatric rehabilitation. Materials and Method: In this cross-sectional comparative content analysis, 24 predefined questions on geriatric rehabilitation were developed from YouTube comments, relevant literature, and clinical experience. Each question was submitted to three large language models under standardized conditions. Anonymized responses were independently evaluated by two experienced physiotherapists for reliability, clinical accuracy, quality, and usefulness, with disagreements resolved by consensus with a specialist physician. Readability was assessed using the Flesch Reading Ease scores. Results: No statistically significant difference was found among the three large language models in terms of reliability scores (p = 0.097). However, significant differences were observed for clinical accuracy, quality, usefulness, readability, and text characteristics (all p = 0.001). ChatGPT-5.2 and DeepSeek V3.2 showed the highest clinical accuracy scores, while ChatGPT-5.2 was superior in terms of quality and usefulness. For readability, ChatGPT-5.2 and Gemini 3 outperformed DeepSeek V3.2. Conclusion: Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance. Nevertheless, because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision. Keywords: Geriatrics; Rehabilitation; Artificial Intelligence; Patient Education as Topic; Health Literacy.

Uğur Sözlü, Selim Mahmut Günay, Sevda Demir Türe et al. · 0 citations
Conference Jul 2026

A Multi-Domain Human Expert Evaluation of Clinical and Behavioral Knowledge in Large Language Models

Large Language Models (LLMs) such as ChatGPT and Gemini are increasingly used to answer medical and psychological questions, yet systematic evaluations across domains with differing reasoning demands remain limited. We present a multi-domain expert evaluation of two state-of-the-art models, ChatGPT Pro (v5.2) and Gemini 3 Pro, across three healthcare domains: Gynecology, Pathology, and Psychology. We curated 300 open-ended, realistic questions, 100 per domain, designed to elicit clinical reasoning, mechanistic interpretation, and conceptual explanation. Responses were independently scored by domain experts using a standardized five-point rubric. Results reveal domain-dependent performance patterns, with both models performing well in guideline-aligned advisory tasks in Gynecology, lower in mechanistic diagnostic contexts in Pathology, and showing divergent strengths in conceptual psychology questions. Quantitative and qualitative analyses highlight recurring limitations in contextual nuance, mechanistic depth, and safety framing. These findings underscore the importance of domainstratified evaluation and expert oversight when deploying LLMs for healthcare information and provide a reproducible framework for future assessments.

Pragna Prahallad, Pranathi Prahallad, Dhrithi P. Desai · 0 citations
Open access Aug 2026

Evaluating Large Language Models in Response to Questions on Substance Use: Helpful or Harmful?

BACKGROUND Individuals with substance use disorders (SUD) are obtaining health-related information from various large language models (LLMs). We aimed to assess whether LLMs provide responses concordant with the current evidence base and whether they provide harmful responses. METHODS Twenty questions related to SUD were posed to three LLMs (Gemini-1.5-pro-001, Claude-3-5-sonnet, and GPT-4) in May 2024. Each response was independently rated by three experienced addiction specialists, and disagreements were resolved by two additional experienced addiction specialists. All raters were blinded to the LLM. Each rater assessed whether (I) a competent addiction specialist would agree with the response, (II) the response contained stigmatizing language as defined by National Institute on Drug Abuse, or (III) the response contained harmful content. RESULTS 88% of responses were rated as competent and 92% as not harmful. Gemini-1.5-pro-001 had the highest rate of competence (95%), followed by Claude-3-5-sonnet and GPT-4 (both 85%). Gemini-1.5-pro-001 produced no harmful responses, while Claude-3-5-sonnet and GPT-4 produced 10% and 15%, respectively. 30% of responses from both Gemini-1.5-pro-001 and Claude-3-5-sonnet had contained stigmatizing language, compared to 10% for GPT-4. CONCLUSIONS While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.

Samuel Maddams, Shan Chen, Danielle S. Bitterman et al. · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists.

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Review Open access Jul 2026

Scalable, context-sensitive psychiatric assessment with large language models and brief diaries

Abstract Background Accurate psychiatric assessment requires understanding a person’s unique experience within their psychosocial context. Clinical interviews have been the gold standard for assessment as the only methods capable of this complex task, but they are time and resource-intensive. Consequently, psychiatric assessment typically relies on patient report surveys that are decontextualized and narrow in scope. This comprehensiveness-scalability tradeoff is a major bottleneck in studying and treating psychopathology. We propose using large language models (LLMs) to score psychopathology from brief personal narratives as a low-burden, context-sensitive solution. Methods Participants (N = 108) completed brief (~1 minute), freeform audio diaries daily for 2 weeks. We used six LLMs to score wide-ranging psychopathology (Internalizing, Detachment, Disinhibition, Antagonism, Anankastia) from the diary transcripts. Leveraging an array of self-report and clinical interview measures, we tested the convergent, discriminant, concurrent, and clinical validity of LLM ratings for between-person differences and within-person fluctuations in psychopathology. Results Supporting convergent and discriminant validity, LLM ratings correlated most strongly with corresponding self-report domains at the between (average convergent r = .42) and within-person (r = .28) levels. LLM and self-report ratings had similar patterns of associations with external variables, except for Anankastia and Antagonism. Further, every LLM-rated domain related to psychopathology ascertained by clinical interview. Conclusions Across multiple forms of validity, we showed that LLMs can assess most major forms of psychopathology from mere minutes of audio. These results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Whitney R. Ringwald, Aman Taxali, Michael Angstadt et al. · 0 citations
Review Open access Jul 2026

Benchmarking large language models against practicing clinicians on psychopathological assessment

Psychiatry’s reliance on language makes LLMs a natural tool for psychopathological assessment, yet structured, item-level assessments from psychiatric clinical interviews remain under-researched. In this proof-of-concept study, 10 LLMs assessed transcripts of three simulated psychiatric interviews across all 100 items of the Association for Methodology and Documentation in Psychiatry (AMDP) system, benchmarked against 108 early-career clinicians rating full audiovisual recordings, using an expert consensus panel as reference. GPT-5.1 and Gemini-3-Pro-Preview achieved the highest accuracy (0.72; 64th percentile of the clinician distribution) using majority voting across three runs with AMDP definitions as context. GPT-5.1, selected for a marginal advantage, showed per-scenario accuracies of 0.81 (depression), 0.76 (mania), and 0.60 (schizophrenia) versus clinician means of 0.79, 0.68, and 0.58. Clinicians and LLMs showed distinct error profiles: clinicians tended to over-infer symptom presence, whereas LLMs more conservatively flagged items as “not assessable” — most pronounced for observation-dependent items but present even for text-assessable items (19.4% vs. 11.4%, p < 0.001). In post hoc simulated disagreement resolutions (2091 clinician pairs; 35.5% disagreements), LLM and board-certified supervision were associated with more accurate resolutions than unsupervised random clinician selection (p < 0.0002). These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.

E. Lenz, J. Naamanka, Wolfgang Trabert et al. · 0 citations

Related blog posts