Skip to content

Author

Martin H. Maurer

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large Language Models.

RATIONALE AND OBJECTIVES Multimodal large language models (LLMs) are increasingly applied to image-based radiology tasks, but their diagnostic accuracy across clinically distinct populations remains poorly characterized. We quantified age-related differences in zero-shot LLM performance for pneumonia detection on pediatric vs. adult chest radiographs and compared generalization gaps with a domain-trained convolutional neural network (CNN) baseline. MATERIALS AND METHODS GPT-5.2 (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Pro (Google) were evaluated zero-shot on balanced pediatric and adult test sets of frontal chest radiographs (n = 1000 each; 500 pneumonia/500 normal). Cohort-specific InceptionV3 CNNs were trained on the remaining development-pool images (pediatric n = 4715; adult n = 13,863) and evaluated on the same test sets. Performance was assessed using the Matthews correlation coefficient (MCC) with 95% bootstrap confidence intervals (CIs); domain shift was quantified as Δ = Adult - Pediatric. RESULTS In pediatrics, the CNN outperformed all LLMs (MCC 0.799, 95% CI 0.766-0.832) vs. GPT-5.2 (0.484, 0.436-0.532), Claude Opus 4.5 (0.470, 0.418-0.521), and Gemini 2.5 Pro (0.272, 0.224-0.316). In adults, all models improved, but the CNN remained best (MCC 0.850, 0.816-0.882). Age-related gains were larger for LLMs (ΔMCC +0.220 [95% CI 0.160-0.281] to +0.466 [0.407-0.525]) than for the CNN (ΔMCC +0.051 [0.005-0.098]), driven mainly by specificity increases. CONCLUSION Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.

Matteo Haupt, Martin H. Maurer · 0 citations
Open access Jul 2026

Benchmarking Multimodal Large Language Models for Cardiopulmonary Findings on Chest Radiographs: Sex-Stratified Discrimination and Operating Characteristics

Commercial MLLMs differ considerably in operating profiles, ranging from ultraconservative to aggressive detection, so that strong aggregate discrimination can mask sensitivity too low for reliable detection.

Matteo Haupt, Arne Bischoff, Myriam Atoubi et al. · 0 citations