Skip to content
Open access

Benchmarking Multimodal Large Language Models for Cardiopulmonary Findings on Chest Radiographs: Sex-Stratified Discrimination and Operating Characteristics

Jul 2026 · Diagnostics · Vol 16 · 0 citations · 32 references
Medicine

TL;DR

Commercial MLLMs differ considerably in operating profiles, ranging from ultraconservative to aggressive detection, so that strong aggregate discrimination can mask sensitivity too low for reliable detection.

Abstract

Background/Objectives: To characterize the zero-shot diagnostic behavior of three commercial multimodal large language models (MLLMs) on cardiopulmonary chest radiograph findings and to assess sex-stratified performance differences. Methods: GPT-5.4, Claude Opus 4.5, and Gemini 2.5 Pro were evaluated in 4500 pathology-specific radiograph evaluations based on frontal chest radiographs from the publicly available CheXpert dataset. Three balanced cohorts of 1500 images each were constructed for cardiomegaly, pulmonary edema, and pleural effusion (375 per sex-by-label subgroup). All models received identical zero-shot prompts requesting binary classification. The primary outcome was area under the receiver operating characteristic curve (AUC-ROC) with 95% bootstrap confidence intervals. Secondary outcomes were sensitivity and specificity. Results: A total of 4500 pathology-specific radiograph evaluations were performed across the three cohorts (2250 male and 2250 female cohort entries; mean age 58.4 ± 18.0 years). GPT-5.4 achieved the highest discrimination (AUC-ROC 0.836–0.883) but showed very low sensitivity (0.043–0.424) with near-perfect specificity (0.977–0.997). Claude Opus 4.5 showed moderate discrimination (AUC-ROC 0.698–0.761) with balanced sensitivity (0.396–0.876) and specificity (0.461–0.863). Gemini 2.5 Pro showed moderate discrimination (AUC-ROC 0.745–0.770) but favored sensitivity (0.673–0.973) at the expense of specificity (0.241–0.804). Sex-stratified analyses showed consistently higher AUC point estimates in male patients for cardiomegaly and pulmonary edema, but smaller and less directional differences for pleural effusion. Conclusions: Commercial MLLMs differ considerably in operating profiles, ranging from ultraconservative to aggressive detection, so that strong aggregate discrimination can mask sensitivity too low for reliable detection. None of the evaluated models are currently suitable for autonomous chest radiograph interpretation. Sex-stratified differences were modest but non-uniform, supporting subgroup-aware reporting rather than reliance on pooled metrics alone.

Read PDF

Similar papers

Aug 2026

Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid Nodules.

BACKGROUND Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear. OBJECTIVE This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules. METHODS This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison. RESULTS All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs. CONCLUSION Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.

Ziman Chen, Yingli Wang, Fei Chen · 0 citations
Open access Jul 2026

Performance evaluation of domain-specific and general-purpose AI models for chest radiograph interpretation: a comparative study.

BACKGROUND Chest radiography remains the most widely used imaging modality worldwide; however, its interpretation is inherently challenging because of overlapping anatomical structures and subtle findings. Recent advances in multimodal large language models (LLMs) have enabled automated radiology report generation, yet their clinical performance relative to domain-specific medical AI systems remains insufficiently validated. OBJECTIVES This study aimed to evaluate the performance and clinical applicability of a domain-specific multimodal AI model (M4CXR) compared with a general-purpose LLM (ChatGPT-4o) for chest radiograph interpretation. METHODS In this retrospective study, 500 anonymized chest radiographs from a single tertiary care center were analyzed. Four board-certified radiologists independently evaluated AI-generated reports from both models. Key outcomes included key finding detection (categorized as complete, partial, or inconsistent), report generation time, and report discrepancies assessed using the RADPEER scoring system. Agreement between original and M4CXR-assisted RADPEER scores was assessed using intraclass correlation coefficients and weighted Cohen's kappa. Statistical analyses included paired t-tests, and chi-square tests. RESULTS M4CXR demonstrated significantly higher report consistency than GPT-4o, with complete concordance observed in 55.8% versus 19.8% of cases, and lower inconsistency rates (25.2% vs. 46.4%, P<.001). The use of M4CXR significantly reduced report generation time compared with unaided interpretation (16.3 ± 12.9 s vs. 179.2 ± 50.4 s, P<.001). RADPEER-based discrepancy analysis revealed no significant differences between original and AI-assisted interpretations. Agreement between original and M4CXR-assisted RADPEER scores showed good reliability (ICC = 0.701), and weighted kappa analysis showed substantial agreement (κw = 0.652). CONCLUSIONS Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions. These findings suggest the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts. Future integration should prioritize human-AI collaboration and prospective multi-center validation.

Tae-Hoon Kim, J. Hong, Jihun Hyun et al. · 0 citations
Open access Aug 2026

Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large Language Models.

RATIONALE AND OBJECTIVES Multimodal large language models (LLMs) are increasingly applied to image-based radiology tasks, but their diagnostic accuracy across clinically distinct populations remains poorly characterized. We quantified age-related differences in zero-shot LLM performance for pneumonia detection on pediatric vs. adult chest radiographs and compared generalization gaps with a domain-trained convolutional neural network (CNN) baseline. MATERIALS AND METHODS GPT-5.2 (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Pro (Google) were evaluated zero-shot on balanced pediatric and adult test sets of frontal chest radiographs (n = 1000 each; 500 pneumonia/500 normal). Cohort-specific InceptionV3 CNNs were trained on the remaining development-pool images (pediatric n = 4715; adult n = 13,863) and evaluated on the same test sets. Performance was assessed using the Matthews correlation coefficient (MCC) with 95% bootstrap confidence intervals (CIs); domain shift was quantified as Δ = Adult - Pediatric. RESULTS In pediatrics, the CNN outperformed all LLMs (MCC 0.799, 95% CI 0.766-0.832) vs. GPT-5.2 (0.484, 0.436-0.532), Claude Opus 4.5 (0.470, 0.418-0.521), and Gemini 2.5 Pro (0.272, 0.224-0.316). In adults, all models improved, but the CNN remained best (MCC 0.850, 0.816-0.882). Age-related gains were larger for LLMs (ΔMCC +0.220 [95% CI 0.160-0.281] to +0.466 [0.407-0.525]) than for the CNN (ΔMCC +0.051 [0.005-0.098]), driven mainly by specificity increases. CONCLUSION Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.

Matteo Haupt, Martin H. Maurer · 0 citations
Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

Abstract Background Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear. Objective This study aimed to evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists. Methods This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from 2 institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (1) basic patient information and imaging findings, (2) scenario A plus chief complaint or clinical history, and (3) scenario B plus key laboratory results. Scenario-based data were input into 3 general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using the McNemar test, and P values were adjusted using the Holm-Bonferroni correction for multiple comparisons. Results A total of 301 patients with pathologically confirmed diseases were included (mean age 53.5, SD 12.0 years; women: n=208, 69.1%). In the liver cohort, a numerical trend toward higher accuracy was observed in scenario C compared with scenario A across all 3 models (scenario C range: 72.3%‐76.2% vs scenario A range: 64.4%‐68.3%); these differences did not reach statistical significance after Holm-Bonferroni correction (all adjusted P>.99). Notably, the DeepSeek-R1 model in scenario C achieved the highest diagnostic accuracy (77/101, 76.2%), with no evidence of a difference compared with radiologists (82/101, 81.2%; adjusted P>.99). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in scenario A for disease diagnosis (68/92, 73.9%), which exceeded its performance in scenario B (64/92, 69.6%) and scenario C (66/92, 71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in scenario B achieved the highest accuracy in this cohort (72/92, 78.3%); however, no statistically significant difference was found compared with radiologists (80/92, 87.0%; adjusted P=.25). In the breast cohort, DeepSeek-R1 achieved the numerically highest diagnostic accuracy in scenario A, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between scenarios A and C (73/108, 67.6% vs 71/108, 65.7%; adjusted P>.99). Conclusions While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations
Open access Aug 2026

Diagnostic Accuracy of ChatGPT Plus (GPT-4o) for the Interpretation of Chest and Extremity Radiographs Against Routine Radiologist Reporting: A Single-Centre, Retrospective, Cross-Sectional Study

Background Multimodal large language models are now widely accessible, but their diagnostic capability on plain-film radiographs is poorly characterized. Most evaluations in radiology address purpose-built convolutional networks rather than general-purpose conversational assistants. Materials and Methods A single-center, retrospective, cross-sectional diagnostic-accuracy study was conducted over 6 months at a tertiary-care teaching hospital in western India. We randomly drew 385 chest and extremity radiographs from PACS, each interpreted independently by ChatGPT Plus (GPT-4o) and compared with the verified radiologist report as the reference standard. Outcomes were sensitivity, specificity, accuracy, likelihood ratios, the diagnostic odds ratio (DOR), Cohen's kappa, and the McNemar exact test, with prespecified subgroup analyses and Wilson 95% confidence intervals. Blinded adjudication of the 39 discordant pairs by an independent consultant radiologist was performed as a sensitivity analysis of the reference standard. Results Among 385 radiographs (228 chest, 157 extremity; abnormal prevalence 25.7%), ChatGPT Plus achieved a sensitivity of 78.8% (95% CI 69.7–85.7), specificity of 93.7% (90.3–96.0), accuracy of 89.9% (86.5–92.5), positive likelihood ratio of 12.5, and a DOR of 55.3. Inter-rater agreement was substantial ( κ  = 0.73; 0.65–0.81), with no systematic discordance (McNemar exact p  = 0.749). Sensitivity was higher for chest than for extremity radiographs (87.3 vs. 63.9%; Fisher's exact p  = 0.010); specificities were comparable. Independent adjudication of the 39 discordant pairs reclassified [X/21] false negatives and [Y/18] false positives as confirmed model errors, [A] as confirmed radiologist omissions or borderline calls, and [B] as legitimately equivocal; the corrected-reference-standard sensitivity and specificity were [S%] and [Sp%] respectively. Conclusion ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs. The model is a plausible supervised educational adjunct or alert application rather than a substitute for expert interpretation; local validation and a human-in-the-loop pathway are prerequisites for any clinical role.

Uzma Khan, Shravani Moholkar, Fatema Kazi et al. · 0 citations
Open access Aug 2026

Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study

The findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.

Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al. · 0 citations