Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.
Aug 2026· Journal of the American Veterinary Medical Association· pp.
1-6
· 0 citations· 20 references
Medicine
Abstract
Objective
To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures.
Methods
In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ.
Results
The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage).
Conclusions
LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios.
Clinical Relevance
Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.
Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures and require specialized pediatric training before clinical implementation, however, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.
U. Kalafat, H. Mutlu, Ramiz Yazıcı et al.· PLoS ONE· 0 citations
Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.
A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al.· Skeletal Radiology· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
The evaluated models lacked the spatial precision and consistency required for the accurate assessment of complex middle ear structures and underscores the necessity for verification by a radiologist and continuous monitoring.
B. Yağcı, Sergen Palaz, E.A. Cetinkaya et al.· Diagnostic and Interventiona...· 0 citations
ChatGPT is currently the most robust model for guideline-based management of UPJO, however, the “deceptive confidence” of models like Copilot poses a risk of misinformation and future integration should explore multimodal capabilities.
İ. Baloğlu, Gökçe Karlı, Ali Emre Çekmece et al.· Pediatric surgery internatio...· 0 citations
Artificial intelligence-based methods demonstrate promising performance in selected image-based TMJ osteoarthritis tasks, but present evidence does not justify autonomous diagnosis or replacement of established clinical and imaging reference standards.
Arturo Arbeláez Ramírez, Daniel Botero Rosas· Diagnostics· 0 citations