Skip to content

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Jul 2026 · Skeletal Radiology · 0 citations · 22 references
Medicine

TL;DR

Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.

View source

Similar papers

Jul 2026

Moderate-to-substantial agreement of ChatGPT-5 for Kellgren–Lawrence grading on synthetic knee radiographs: a controlled cross-sectional observer agreement study

ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions under curated synthetic dataset, however, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.

Öner Kılınç, Elif Altunel Kılınç, N. Çabuk Çelik · 0 citations
Open access Aug 2026

Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

M. Hamdan, A. Harati, A. Al-bakheet et al. · 0 citations
Open access Aug 2026

Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study

The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.

Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al. · 0 citations
Aug 2026

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

S. Okur, Ç. Özkalıpçı, Büşra Baykal et al. · 0 citations