Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.
ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions under curated synthetic dataset, however, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.
Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
M. Hamdan, A. Harati, A. Al-bakheet et al.· medRxiv· 0 citations
The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.
Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al.· Journal of Medical Internet...· 0 citations
Objective
To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures.
Methods
In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ.
Results
The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage).
Conclusions
LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios.
Clinical Relevance
Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.
S. Okur, Ç. Özkalıpçı, Büşra Baykal et al.· Journal of the American Vete...· 0 citations