Aug 2026· npj Digital Medicine· Vol 9· 0 citations· 48 references
Medicine
TL;DR
GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.
Abstract
Multidisciplinary tumour boards (MDTs) are the standard for gastrointestinal oncological decision-making but remain resource-intensive. Whether specialty-specific role prompting induces genuinely distinct clinical reasoning in large language models (LLMs)—or merely role-appropriate language around an invariant output—has not been systematically tested. We applied five zero-shot prompting frameworks and a majority-vote ensemble to GPT-5 across 100 gastrointestinal oncology cases with MDT-validated decisions: a simulated MDT, multi-expert deliberation, three specialist personas, and a majority-vote ensemble. Concordance with MDT recommendations ranged from 78% to 87%, with no significant inter-framework differences (Cochran’s Q = 8.46, p = 0.133). Specialty-characteristic language was near-universal (97–100%) but uncorrelated with accuracy. Embedding analysis revealed high semantic similarity across personas (cosine similarity 0.805–0.836; η² = 0.049), contrasting with substantially greater output separation under multi-expert deliberation (η² = 0.554–0.581). GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark, however, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.
Qian Li, Yongxin Li, Chao Ye et al.· Frontiers in Artificial Inte...· 0 citations
Clinical practice guidelines require expert synthesis that large language models (LLMs) might partly automate, yet their ability to reproduce clinically actionable recommendations is poorly quantified. We evaluate an LLM (Claude Sonnet 4.6) against the 103 recommendations of the Spanish enhanced-recovery guideline Vía RICA 2026, grouped in 17 bundles. The model used the panel’s own closed corpus (617 documents) in a multilingual retrievalaugmented generation pipeline. Concordance was assessed twice: by optimal 1:1 bipartite matching (Hungarian) on cosine similarity, and by an LLM-as-a-judge clinical adjudicator (Claude Haiku 4.5) validated against a three-clinician panel (Fleiss’ κ = 0.538). The two schemes bracket a micro F1 of 0.61–0.69 and reveal four findings: (i) a systematic granularity bias, producing 1–8 recommendations per bundle regardless of ground-truth size; (ii) failure of cosine similarity to discriminate within narrow clinical domains; (iii) high reference-concordance precision (0.70–0.81) despite low exhaustiveness; and (iv) no transfer of the GRADE fields, evidence level agreeing no better than chance and strength systematically downgraded. An eight-fold larger retrieval budget left it intact. A corpus audit found 25 documents that formulate recommendations; excluding them lowers judged micro F1 to 0.602. The results delimit the current utility of generative AI for guideline development.
Andrea Moral, Antonio Arroyo, Juan Aparicio et al.· Machine Learning and Knowled...· 0 citations
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.
Michael Fiorino, Mei Hainline, Tanay Poddar et al.· Neuro-Oncology Advances· 0 citations