Skip to content

Author

Dingyu Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

The illusion of clinical reasoning: a benchmark reveals the pervasive gap in vision-language models for clinical competency

The rapid integration of foundation models into clinical practice and their use for public health inquiries necessitates a rigorous evaluation of their true clinical reasoning capabilities, which extends beyond success on narrow examinations. Current benchmarks, often based on medical licensing exams or curated vignettes, fail to capture the integrated, multimodal reasoning required in real-world patient care. To address this gap, we developed the Bones and Joints (B&J) Benchmark, a comprehensive evaluation framework comprising 1245 questions derived from real-world patient cases in orthopedics and sports medicine. This benchmark assesses models across seven core tasks that mirror the clinical reasoning pathway, including knowledge recall, text interpretation, image interpretation, diagnosis generation, treatment planning, and the underlying rationale. We evaluated 14 vision-language models (VLMs) and six large language models (LLMs), comparing their performance against expert-derived ground truth. Our findings reveal a pronounced performance gap. While state-of-the-art models achieved high accuracy, exceeding 90% on structured multiple-choice questions, their performance markedly declined on open-ended tasks requiring multimodal integration, with accuracy scarcely reaching 60%. VLMs demonstrated substantial limitations in interpreting medical images and frequently exhibited text-driven hallucinations. Notably, medical-specific models showed no consistent advantage over general-purpose counterparts. These results indicate that current foundation models face significant challenges in achieving independent clinical competence within highly specialized musculoskeletal fields. Their safe deployment should be limited to supportive, text-based roles, while advancement in core clinical tasks awaits fundamental breakthroughs in multimodal integration and visual understanding.

Dingyu Wang, Z. L. Yuan, Jiajun Liu et al. · 0 citations