This study provides a framework for evaluating LLMs in psychiatry and supporting their safe, evidence-based integration into mental health practice and is anticipated to improve diagnostic accuracy, especially among non-specialists.
Abstract
Background
Psychiatric medicine presents unique diagnostic and therapeutic challenges, often involving multimorbidity, polypharmacy, and atypical presentations requiring complex reasoning. Artificial Intelligence, particularly Large Language Models (LLMs), is emerging as a support tool in these settings. However, the clinical validity, interpretability, and reliability of LLMs in psychiatry remain largely unexplored, particularly their ability to generate transparent, guideline-consistent reasoning.
Aims
This study evaluates the clinical reasoning capabilities of LLMs in complex psychiatric scenarios. The primary aim is to assess the validity of their long chain-of-thought (CoT) reasoning. A secondary aim is to determine whether LLM assistance improves clinicians’ diagnostic accuracy.
Method
Three LLMs, Gemini 2.5, Grok, and DeepSeek R1, will be assessed using ten complex psychiatric cases sourced from a non-public clinical manual and rated with the Amsterdam Clinical Challenge Scale. Each model's CoT response will be evaluated by blinded panels of psychiatrists, residents, and general practitioners using standardized metrics for factual accuracy, coherence, and medical plausibility. In a second phase, clinicians will answer thirty diagnostic questions with and without support from the best-performing LLM. The study uses step-by-step reasoning prompts and few-shot examples to elicit detailed responses and includes bias mitigation strategies such as randomization, blinding, and statistical controls.
Results
Analyses will assess inter-rater reliability, metric redundancy, and reasoning quality. Closed-source models are expected to outperform open-source ones. LLM assistance is anticipated to improve diagnostic accuracy, especially among non-specialists.
Conclusions
This study provides a framework for evaluating LLMs in psychiatry and supporting their safe, evidence-based integration into mental health practice.
Abstract Background Despite the availability of more than 25 licensed antipsychotics worldwide, prescribing decisions remain constrained by oversimplified classifications and fragmented guideline recommendations. The conventional dichotomy between “typical” and “atypical” antipsychotics fails to capture meaningful diff...
R. McCutcheon· International Journal of Neu...· 0 citations
Abstract Background Clinicians across disciplines use the Diagnostic and Statistical Manual of Mental Disorders (DSM) to describe, diagnose, and plan treatment for mental health conditions. This means that future iterations of the DSM need to incorporate information that captures diverse expressions of symptomatology a...
D. Clarke, L. Yousif· International Journal of Neu...· 0 citations
Abstract Background The Diagnostic and Statistical Manual of Mental Disorders (DSM) is widely used by healthcare providers. It is a common way for clinicians and researchers across disciplines to understand and codify mental and substance use disorders. In 2024, the American Psychiatric Association’s Board of Trustees...
J. Alpert· International Journal of Neu...· 0 citations
Persistent comorbidity, diagnostic instability and the scarcity of disorder-specific causes have encouraged the proposal that a single superordinate dimension underlies liability to all common mental disorders. This proposition, expressed empirically as the general factor of psychopathology and framed conceptually as a...
Luiz Costa, Fernando Martins Castanheira, Giovanna Azevedo Rodrigues et al.· International Neuropsychiatr...· 1 citation
"AI psychosis"has entered public and clinical discourse as a label for the onset or exacerbation of psychotic symptoms, most commonly delusions, following intensive interaction with large language model (LLM)-based chatbots. Current evidence is limited to media reports, case reports, and early observational data, yet t...
Joshua Au Yeung, H. Morrin, Vincent Ng et al.· 0 citations
This short communication explores the utility of generative artificial intelligence (GenAI) systems in helping students prepare for clinical examinations. Four clinical scenarios from an objective structured clinical examination (OSCE) in psychiatry were presented to GPT-4o mini. This GenAI system was asked to respond...
M. Henning, Zhao-Chu Geng, Christian U. Krägeloh et al.· Education in Medicine Journa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.