Skip to content

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Aug 2026 · Journal of the American Veterinary Medical Association · pp. 1-6 · 0 citations · 20 references
Medicine

Abstract

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

View source

Similar papers

Open access Jul 2026

Artificial intelligence meets pediatric orthopedics: A comparative analysis of ChatGPT-4o, Gemini 2.0, and Claude 3.5 in detecting supracondylar humeral fractures

Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures and require specialized pediatric training before clinical implementation, however, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.

U. Kalafat, H. Mutlu, Ramiz Yazıcı et al. · 0 citations
Jul 2026

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.

A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al. · 0 citations
Open access Aug 2026

Diagnostic performance and short-interval reproducibility of multimodal large language models in differentiating cholesteatoma from chronic otitis media using key-image temporal bone high-resolution computed tomography.

The evaluated models lacked the spatial precision and consistency required for the accurate assessment of complex middle ear structures and underscores the necessity for verification by a radiologist and continuous monitoring.

B. Yağcı, Sergen Palaz, E.A. Cetinkaya et al. · 0 citations
Jul 2026

Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios

ChatGPT is currently the most robust model for guideline-based management of UPJO, however, the “deceptive confidence” of models like Copilot poses a risk of misinformation and future integration should explore multimodal capabilities.

İ. Baloğlu, Gökçe Karlı, Ali Emre Çekmece et al. · 0 citations
Review Open access Aug 2026

Artificial Intelligence for Diagnosis of Temporomandibular and Cranio-Cervico-Mandibular Musculoskeletal Disorders: A Systematic Review and Exploratory Diagnostic Test Accuracy Meta-Analysis

Artificial intelligence-based methods demonstrate promising performance in selected image-based TMJ osteoarthritis tasks, but present evidence does not justify autonomous diagnosis or replacement of established clinical and imaging reference standards.

Arturo Arbeláez Ramírez, Daniel Botero Rosas · 0 citations