Assessing the Authenticity of Artificial Intelligence Appraisals of Clinical Psychiatric Scenarios: A Preliminary Investigation
Abstract
This short communication explores the utility of generative artificial intelligence (GenAI) systems in helping students prepare for clinical examinations. Four clinical scenarios from an objective structured clinical examination (OSCE) in psychiatry were presented to GPT-4o mini. This GenAI system was asked to respond to four questions related to the clinical psychiatric scenarios. Two questions were designed to be straightforward and contextually clear, whereas two incorporated irrelevant information. The GenAI responses demonstrated noticeable difficulty in accurately and authentically interpreting the scenarios, particularly when extraneous content was introduced. While the answers to the first two questions were reasonably accurate, the responses to the remaining two were unacceptable, highlighting the system’s inability to decipher ambiguous or unclear input. Currently, GenAI systems appear to perform well at solving clinical cases when the data are straightforward and rule-based. However, GPT-4o mini demonstrated difficulty in interpreting scenarios complicated by emotional nuances or distracting elements.