Skip to content

Author

M. Hirschmann

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.

PURPOSE To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA). METHODS This is a cross-sectional comparative study. Twenty questions on TKA supported by moderate-strong level of evidence were randomly selected from the WEMA. Three orthopaedic surgeons independently performed a blinded assessment of all LLM-generated responses. An adapted version of the QUEST rating system, a comprehensive framework designed for the objective human assessment of LLM performance across healthcare-related subdomains, was used. Furthermore, the same three evaluators subjectively selected the best-performing LLM response for each question. RESULTS The four LLMs presented statistically significant differences in overall performance based on the QUEST framework (score range 1-5): Gemini 2.5; 4.77 ± 0.06, Claude 4; 4.72 ± 0.07, ChatGPT-5; 4.70 ± 0.08 and GROK 4; 4.62 ± 0.09 (p < 0.001). Gemini 2.5 achieved the highest scores in the Accuracy (4.58 ± 0.70), Comprehensiveness (4.87 ± 0.34) and Trust (4.52 ± 0.62) dimensions. However, Claude 4 obtained the highest score for the Currency (4.23 ± 0.67) dimension. When assessors subjectively selected the superior answer for each question, Claude 4 was chosen most frequently, in 46.7% of cases, followed by ChatGPT-5 in 25.4%, Gemini 2.5 in 22.9% and GROK 4 in 7.5% of cases. CONCLUSIONS Four general-purpose LLMs (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) demonstrated good overall performance when addressing specialised clinical questions related to TKA. No single model consistently outperformed the others across all evaluated domains. LEVEL OF EVIDENCE Level V.

Oriol Pujol, R. Ferrer, Alex Coelho et al. · 0 citations