Skip to content

Author

Shravani Moholkar

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Diagnostic Accuracy of ChatGPT Plus (GPT-4o) for the Interpretation of Chest and Extremity Radiographs Against Routine Radiologist Reporting: A Single-Centre, Retrospective, Cross-Sectional Study

Background Multimodal large language models are now widely accessible, but their diagnostic capability on plain-film radiographs is poorly characterized. Most evaluations in radiology address purpose-built convolutional networks rather than general-purpose conversational assistants. Materials and Methods A single-center, retrospective, cross-sectional diagnostic-accuracy study was conducted over 6 months at a tertiary-care teaching hospital in western India. We randomly drew 385 chest and extremity radiographs from PACS, each interpreted independently by ChatGPT Plus (GPT-4o) and compared with the verified radiologist report as the reference standard. Outcomes were sensitivity, specificity, accuracy, likelihood ratios, the diagnostic odds ratio (DOR), Cohen's kappa, and the McNemar exact test, with prespecified subgroup analyses and Wilson 95% confidence intervals. Blinded adjudication of the 39 discordant pairs by an independent consultant radiologist was performed as a sensitivity analysis of the reference standard. Results Among 385 radiographs (228 chest, 157 extremity; abnormal prevalence 25.7%), ChatGPT Plus achieved a sensitivity of 78.8% (95% CI 69.7–85.7), specificity of 93.7% (90.3–96.0), accuracy of 89.9% (86.5–92.5), positive likelihood ratio of 12.5, and a DOR of 55.3. Inter-rater agreement was substantial ( κ  = 0.73; 0.65–0.81), with no systematic discordance (McNemar exact p  = 0.749). Sensitivity was higher for chest than for extremity radiographs (87.3 vs. 63.9%; Fisher's exact p  = 0.010); specificities were comparable. Independent adjudication of the 39 discordant pairs reclassified [X/21] false negatives and [Y/18] false positives as confirmed model errors, [A] as confirmed radiologist omissions or borderline calls, and [B] as legitimately equivocal; the corrected-reference-standard sensitivity and specificity were [S%] and [Sp%] respectively. Conclusion ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs. The model is a plausible supervised educational adjunct or alert application rather than a substitute for expert interpretation; local validation and a human-in-the-loop pathway are prerequisites for any clinical role.

Uzma Khan, Shravani Moholkar, Fatema Kazi et al. · 0 citations