Skip to content
Open access

Evaluating Large Language Models for AI-Assisted Decision Support in Legal Capacity Assessment: A Comparative Study Using Interdisciplinary Medical Board Recommendations as the Expert Medical Reference Standard

Jul 2026 · Healthcare · Vol 14 · 0 citations · 53 references
Medicine

TL;DR

LLMs demonstrated agreement with IMBD recommendations on standardized medico-legal case vignettes, supporting further investigation of their potential role as AI-assisted decision-support tools under expert supervision.

Abstract

Background: Legal capacity assessment requires multidisciplinary evaluation integrating cognitive, functional, neurological, psychiatric, and medico-legal information. Although large language models (LLMs) have shown promise in structured clinical reasoning, their role in supporting legal capacity assessment remains unclear. This study evaluated the performance of LLMs as AI-assisted decision-support tools using standardized medico-legal case vignettes, with interdisciplinary medical board (IMBD) recommendations under the Turkish Civil Code (TCC) as the expert medical reference standard; IMBD recommendations constitute an expert medical reference standard rather than final judicial determinations. Methods: We retrospectively analyzed 234 court-referred adult cases (2018–2024). Standardized, anonymized medico-legal case vignettes were independently evaluated by ChatGPT-5.2, Gemini 3 Pro, and Claude 4.5 Sonnet. Model outputs were compared with IMBD recommendations. The models were evaluated as AI-assisted decision-support tools and did not replace or influence clinical or judicial decision-making. Performance was assessed using accuracy, macro-F1, Cohen’s κ, AUC, calibration, decision-curve analysis, test–retest reliability, and human-factor outcomes; probability-based metrics were derived from secondary logistic models fitted to the categorical model outputs. Results: Article 405 was the most frequent outcome (65.8%). Dementia increased the likelihood of Article 405 recommendations (OR 6.5, 95% CI 1.2–35.2; p = 0.029), whereas higher Activities of Daily Living scores were protective (OR 0.96; p = 0.004). Gemini 3 Pro achieved the highest accuracy (86.3%), while ChatGPT-5.2 achieved the highest macro-F1 score (0.79) and AUC (0.91). Agreement with IMBD recommendations ranged from κ = 0.60 to 0.73, with high temporal stability (κ = 0.87–0.95); pairwise differences between the three models were not statistically significant after Holm correction. Performance was highest for cases in which Article 405 was recommended and for cases in which no guardianship was recommended, but remained limited for Article 408 (29.6% accuracy). The mean System Usability Scale score was 72.4, and safety flags were identified in 4 of 234 cases (1.7% of cases, corresponding to 4 of 702 individual model outputs). Conclusions: LLMs demonstrated agreement with IMBD recommendations on standardized medico-legal case vignettes, supporting further investigation of their potential role as AI-assisted decision-support tools under expert supervision. Because errors in this domain can directly affect fundamental rights, the use of AI in the legal system and in sensitive medical fields carries substantial risks and must remain strictly limited to expert-supervised decision support. Further prospective studies are needed to evaluate their safe integration into medico-legal practice.

Read PDF

Similar papers

Review Open access Aug 2026

A Human-Governed Clinical Informatics Framework for Safe AI-Assisted Mental Health Counseling: Secondary Framework Development and Requirement Mapping Study.

BACKGROUND Natural language processing and large language model systems are increasingly used to support mental health documentation, screening, and follow-up planning. In counseling contexts, model outputs may influence diagnostic framing, risk recognition, and clinical record content. Static performance metrics and fluent generated summaries are not sufficient to support safe implementation without governance, safety gating, human review, and monitoring. OBJECTIVE This study aimed to develop a human-governed clinical informatics framework for safe AI-assisted mental health counseling and make the formative evidence base and requirement-mapping process traceable. METHODS We conducted a secondary framework development and requirement mapping study using the Korean AI Hub psychological counseling dataset, official data description and use documents, released KLUE-BERT risk prediction model materials, released KoAlpaca summary generation resources, and a deidentified 139-case rule-based summary safety screening audit table derived from the original summary comparison file. Raw counseling transcript text, reference summary full text, and generated summary full text are not included in the manuscript or supplementary materials. We extracted failure modes from documented data and model characteristics, released code and configuration files, documentation-reported model metrics, and rule-based proxy flags. Each failure mode was mapped to safety controls, operational criteria, and deployment-level requirements. RESULTS The official documents described 1661 counseling sessions and 465,474 paragraph-level tokens across depression, anxiety disorder, addiction, and normal control groups. Of the 1661 sessions, the documented split included 1339 (80.6%) training, 173 (10.4%) validation, and 149 (9%) test sessions. The summary generation materials documented 1278 training summaries and 139 test summaries. Documentation-reported model metrics included KLUE-BERT accuracies of 71.43% for depression, 73.53% for anxiety, and 66.67% for addiction and KoAlpaca BERTScore precision, recall, and F1-score values of 62.13%, 59.56%, and 60.80%, respectively. The 139-case screening table contained 77 (55.4%) depression, 31 (22.3%) anxiety, and 31 (22.3%) addiction cases. Rule trigger rates included unsupported content proxy flags in 41% (57/139) of cases, overdiagnostic expression proxy flags in 31.7% (44/139) of cases, medicalized expression proxy flags in 54.7% (76/139) of cases, and any rule-based proxy flag in 91.4% (127/139) of cases. These values are conservative rule trigger rates rather than confirmed clinical error rates. The findings informed a 7-stage workflow, 6 safety control layers, an operational safety gate, a workflow-to-control crosswalk, deployment-level transition criteria, and a constructed high-risk example. CONCLUSIONS AI-assisted mental health counseling should be implemented as a governed clinical information workflow rather than as an autonomous diagnostic or documentation pathway. The proposed framework specifies safeguards and validation requirements for future supervised evaluations, but it does not itself establish clinical safety or clinical effectiveness. Prospective simulation, clinician usability testing, patient or client feedback, and independent expert validation remain necessary before routine deployment.

Mi-Ae Yang, Kang-Su Ha · 0 citations
Review Open access Aug 2026

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.

Sanjay Basu, Sadiq Y. Patel, Parth Sheth et al. · 0 citations
Open access Jul 2026

Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study

Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool.

Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al. · 0 citations
Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations
Open access Apr 2026

Should I and Can I Use AI for Forensic Psychiatry Report Writing?

Abstract The rapid evolution of AI, particularly large language models (LLMs), has renewed interest in their potential role in forensic psychiatry report writing. Recent evidence demonstrates that contemporary LLMs perform well in selected medical knowledge, documentation, and information management tasks and may reduce the administrative burden when deployed under appropriate clinical supervision. However, forensic psychiatric reports differ fundamentally from routine clinical documentation. They constitute expert evidence prepared for legal proceedings and therefore require transparent reasoning, explicit weighing of competing evidence, a robust factual foundation, and personal professional accountability. This viewpoint examines whether AI can and should be used in forensic psychiatry report writing by integrating recent empirical evidence, forensic psychiatry guidance, legal and regulatory frameworks, and emerging governance recommendations. Rather than comparing AI with an idealized human evaluator, the manuscript argues that the appropriate comparison is between 2 imperfect systems of reasoning. Human experts remain susceptible to cognitive biases, omission errors, and disagreement, whereas contemporary LLMs exhibit distinct vulnerabilities, including hallucinations, hidden omissions, probabilistic reasoning, and limited explainability. Although the mechanisms differ, both may ultimately compromise the reliability of expert evidence if left unchecked. Current evidence supports AI for bounded, reversible, and independently verifiable tasks, such as document organization, chronology construction, indexing, transcription, and structured summarization, particularly within secure and validated environments. By contrast, there remains insufficient evidence to support AI-assisted generation or material shaping of psycholegal reasoning, credibility assessments, or final forensic opinions. Because these activities require interpretation, accountability, and reasoning that can withstand judicial scrutiny, they remain fundamentally human responsibilities. The most defensible implementation model is, therefore, one of AI around the report rather than AI writing the report, in which AI serves as a supervised productivity tool while the forensic psychiatrist retains full authorship, accountability, and justification of all substantive conclusions.

Alexandre Hudon · 0 citations