Skip to content

Author

Tingting Hu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Aug 2026

Multimodal large language model–assisted workflow for image-dependent very short answer question generation and response grading in neurosurgical residency assessment: development and validation study

Very short answer questions (VSAQs) reduce cueing but require substantial specialist effort to develop and grade, particularly when they involve neuroimaging. We evaluated an expert-supervised workflow using multimodal large language models (LLMs) to generate image-dependent neurosurgical VSAQs and support grading of resident responses. In this three-phase, single-center study, two multimodal LLMs generated one VSAQ from each of 36 deidentified neurosurgical cases using identical clinical information and three MRI images. Three senior neurosurgeons rated the 72 candidate items for scientific accuracy, image dependency, educational value, and clarity. The higher-scoring item for each case formed a 36-item Golden Set administered to 10 neurosurgery residents. Two senior neurosurgeons independently scored all 360 responses using a 0–2 rubric and resolved disagreements by consensus. ChatGPT graded the same responses using case-specific structured inputs and a fixed prompt. Between-model ratings were compared using paired t tests; agreement was assessed using exact agreement and quadratic weighted kappa, and participant-level totals using Pearson correlation. Both models received high ratings for scientific accuracy, educational value, and clarity, with no significant between-model differences. ChatGPT had higher image dependency ratings than Gemini (mean 4.06, SD 0.44 vs. 3.58, SD 0.52; mean difference 0.48, 95% CI 0.26–0.71; P < 0.001), whereas total scores did not differ significantly (mean 16.34, SD 0.86 vs. 15.93, SD 1.12; P = 0.115). Across 360 resident responses, 275 (76.4%) were fully correct, 28 (7.8%) partially correct, and 57 (15.8%) incorrect; total scores ranged from 47 to 63 out of 72. Before consensus, the experts agreed on 347 responses (96.4%; quadratic weighted kappa = 0.967). AI grading agreed exactly with the final expert scores for 324 responses (90.0%; quadratic weighted kappa = 0.905), and participant-level totals were strongly correlated (Pearson r = 0.970). Under the structured conditions evaluated, multimodal LLMs showed preliminary feasibility for expert-supervised drafting and first-pass grading of image-dependent neurosurgical VSAQs. Larger multicenter studies and external validation are needed before routine implementation or use in high-stakes assessment.

Ye Yuan, Baoping Zheng, Chengye Yao et al. · 0 citations