Skip to content

Moderate-to-substantial agreement of ChatGPT-5 for Kellgren–Lawrence grading on synthetic knee radiographs: a controlled cross-sectional observer agreement study

Jul 2026 · Rheumatology International · Vol 46 · 0 citations · 41 references
Medicine

TL;DR

ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions under curated synthetic dataset, however, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.

View source

Similar papers

Jul 2026

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.

A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al. · 0 citations
Open access Jul 2026

Assessment of Large Language Models and Expert Clinicians in Grading Mandibular Third Molar Extraction Difficulty

Background The Pederson difficulty score (PDS) is widely used to assess mandibular third molar extraction difficulty, but is subject to interobserver variability because no universally accepted reference standard exists. This study investigates whether multimodal large language models (LLMs) demonstrate systematic grading tendencies comparable to those of human clinicians in the absence of a definitive standard. Methods This retrospective cross-sectional study included 100 panoramic radiographs. Two LLMs (GPT and Gemini) and two oral and maxillofacial surgeons independently graded extraction difficulty using the PDS. No reference standard or gold-standard rater was designated; all raters were treated as independent assessors. Agreement was assessed using quadratic-weighted kappa and intraclass correlation coefficient values. Systematic bias was analysed using Bland–Altman plots and ordinal logistic regression. The effects of prompt language and session protocol were also evaluated. Results The interexpert agreement was substantial (κ = 0.754). GPT demonstrated moderate agreement with experts (κ = 0.564-0.590), whereas Gemini exhibited fair-to-moderate agreement (κ = 0.356-0.461). Overall, the four-rater reliability was moderate (intraclass correlation coefficient = 0.553). Both LLMs displayed minimal systematic bias (GPT: −0.13; Gemini: −0.08), with no significant differences in grading tendencies compared with human experts (P > .05). However, sequential-session evaluation significantly reduced the scoring consistency for both models (GPT: P = .001; Gemini: P < .001), whereas prompt language had no significant effect on total PDS scores (GPT: P = .543; Gemini: P = .386). A component-level linguistic variation was observed only in GPT’s assessment of the ramus relationship (P = .003). Conclusion The examined LLMs showed grading tendencies partly comparable to those of human clinicians, without evidence of systematic directional bias; however, agreement remained below interexpert levels and was sensitive to conversational context. These findings should not be interpreted as evidence of diagnostic accuracy or clinical reliability, as no reference standard was established. Clinical relevance LLMs may provide supplementary grading information if used with independent-session protocols; however, their role in clinical decision-making requires further validation against defined outcome standards.

S. Yoon, Jeong-Hun Yoo, Taejun Kim et al. · 0 citations
Review Open access Aug 2026

Inter- and Intraobserver Reliability of Wrist Arthroscopy for TFCC Tears: An International Study.

PURPOSE To evaluate the interobserver and intraobserver reliability of wrist arthroscopy for the diagnosis and classification of triangular fibrocartilage complex (TFCC) tears among experienced European hand surgeons, to assess the effect of a consensus meeting on reliability, and to determine whether reliability differs between different tears. METHODS Five hand surgeons from Europe reviewed 43 standardized wrist arthroscopy cases at two time points separated by 6 weeks. Each case included a clinical vignette, radiographs, and arthroscopy video. Surgeons assessed for the presence of a TFCC tear, classified tears using the Palmer and Atzei-Luchetti systems, and recorded treatment decisions. A consensus meeting was held between rounds. For the whole cohort, interobserver reliability was assessed using Fleiss' kappa and intraobserver reliability using Cohen's kappa. For subgroup analyses, interobserver reliability was assessed using pairwise agreement. RESULTS Interobserver agreement for TFCC tear presence was fair in both rounds, with modest improvement after the consensus meeting. Agreement for TFCC classification was slight-to-fair, and agreement for treatment decisions was low. Intraobserver agreement was consistently higher than interobserver agreement and generally moderate. Subgroup analyses showed higher interobserver agreement for central tears than peripheral tears in assessing TFCC tear presence/absence (central: 94.3% in both rounds versus peripheral: 80.0% and 85.0%). In contrast, for Palmer classification, agreement was higher for peripheral tears, particularly in Round 2 (peripheral: 61.3% versus central: 38.6%). Similar patterns were observed when combined tears were included. The consensus meeting did not meaningfully increase agreement within subgroups between rounds. CONCLUSIONS Arthroscopy demonstrated only moderate reliability for the assessment of TFCC pathology in this study, even among experienced hand surgeons. Although tear presence is commonly recognized, substantial variability persists in classification and decision making for treatment. CLINICAL RELEVANCE These findings highlight limitations of current arthroscopic classification systems and underscore the need for improved education and classification optimization to enhance diagnostic reproducibility.

M. Räisänen, Francesco Smeraglia, Martin Clementson et al. · 0 citations
Review Open access Aug 2026

SP 9.02 Robustness of Retrospective Surgical Complication Grading for Time-Series Data Capture: An Inter-Observer Agreement Analysis Using Clavien-Dindo Classification

It is demonstrated that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding, however, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days).

Sam Pathmanathan, Shi Lam, Sarah Alsaad et al. · 0 citations
Aug 2026

Structured ACR Bone-RADS Versus Unstructured Assessment of Bone Tumors on Radiographs: A Multicenter, Multireader Study Across Experience Levels.

PURPOSE To compare the diagnostic performance and interreader agreement of structured ACR Bone Reporting and Data System (Bone-RADS) versus unstructured assessment of bone tumors on radiographs and to assess their effects on management recommendations. MATERIALS AND METHODS Multicenter, multireader study with a primary retrospective cohort (n = 1,423; 4 centers), external retrospective cohort (n = 354; 6 centers), and prospective cohort (n = 152; 2 of these 10 centers; written informed consent obtained from all prospective participants). After a standardized training session with practice cases, nine musculoskeletal radiologists (early career 3-5 years; midcareer 15-20 years; late career >20 years) independently performed unstructured classification and Bone-RADS 1 to 4 grading with a standardized 4-week washout. Primary end point was reader-averaged difference in area under the curve (ΔAUC) via an Obuchowski-Rockette-Hillis multireader, multicase receiver operating characteristic model; interreader agreement and management recommendation reclassification were also assessed. RESULTS Overall reader-averaged AUC did not differ significantly between unstructured assessment and Bone-RADS across the primary retrospective, external retrospective, and prospective cohorts (ΔAUC: -0.0012, 0.0160, and 0.0053, respectively; all P > .05). However, early-career readers demonstrated significant AUC improvements with Bone-RADS in the primary and external cohorts ΔAUC: +0.0302 and +0.0615; both P < .001), with a similar but nonsignificant difference in the prospective cohort (ΔAUC = +0.0201; P = .572). In contrast, midcareer readers showed nonsignificant changes, and late-career readers exhibited slight performance declines. Interreader agreement increased across all cohorts (eg, in the prospective cohort, early-career Fleiss κ improved from 0.394 to 0.502). Feature-level agreement was highest for pathologic fracture but lower for margins and endosteal erosion (κ ∼0.518-0.660). Malignancy rates increased monotonically across Bone-RADS categories 1 to 4 (9.44%, 25.44%, 46.67%, and 79.67%, respectively). Among early-career readers, Bone-RADS shifted management recommendations by increasing referral or further workup recommendations for potentially malignant lesions in the primary (+8.61%), external (+15.30%), and prospective (+7.14%) cohorts, while concurrently reducing routine surveillance recommendation rates for benign lesions (-17.27%, -15.90%, and -19.51%, respectively). CONCLUSION Bone-RADS provides stable risk stratification, improves diagnostic performance and interreader agreement in early-career readers, and shifts management recommendations toward greater referral or workup of potentially malignant lesions. Its sensitivity-oriented design may increase overmanagement of some benign lesions. To mitigate this, future refinements should consider incorporating age and anatomic-site information, adopting context-aware management thresholds, and providing structured training on lower-agreement features such as margin classification and endosteal erosion.

Chunlin Song, Tianzi Jiang, Tongyu Wang et al. · 0 citations
Open access Aug 2026

Inter-Rater Reliability And Agreement Of The Diagnostic Assessment Scale For Kushtha In Papulosquamous Skin Disease: A Cross-Sectional Two-Rater Study

Background and objectives: The Diagnostic Assessment Scale for Kushtha (DASK) is a newly developed 26-item ordinal instrument that renders the classical threefold Ayurvedic examination as a severity score and a dosha attribution. No reliability estimate has been published. This study estimated its inter-rater reliability, agreement, and the measurement error attaching to an individual score. Methods: Sixty consecutive patients with consultant-confirmed papulosquamous disease (43 psoriasis, 17 lichen planus) were each assessed independently by two postgraduate-qualified Ayurvedic physicians in a single clinical session, giving a fully crossed design of 3,120 item scores. Reliability was estimated as the intraclass correlation coefficient ICC(2,1); agreement as the standard error of measurement (SEM), minimal detectable change (MDC95) and Bland–Altman limits; and item-level agreement as quadratic weighted kappa with Gwet’s AC2. Internal consistency and categorical agreement were also assessed, following the GRRAS guidelines. Results: No data were missing. ICC(2,1) was 0.952 (95% confidence interval 0.92–0.97) for the total score and 0.945, 0.920 and 0.920 for the Vata, Pitta and Kapha subscales. Between-patient variance accounted for 95.2% of total variance and the rater effect for 0.0–0.6%. SEM was 2.60 points and MDC95 7.21 points on the 0–104 scale, and all 26 items reached substantial or almost-perfect agreement (weighted kappa 0.732–0.963). The categorical outputs were less reliable than the scores generating them: severity band kappa 0.800 and dosha attribution kappa 0.640 (0.37–0.85), every disagreement occurring at a cut-point or a narrow margin. Kapha was reproduced consistently yet was not internally coherent (alpha 0.43 and 0.42). Conclusion: DASK measures the severity of papulosquamous Kushtha reliably enough for group comparison and, against a threshold of 8 points, for individual monitoring. Agreement on its categorical outputs was lower than on the continuous scores from which they are derived, and the Kapha subscale was reproduced consistently between raters without cohering internally.

A. M, N. S, Saritha T · 0 citations