How Often Do Large Language Models Agree with Each Other—And with the Truth? A Consensus- and Complexity-Stratified Analysis of Data Extraction for Neuroimaging AI
Aug 2026· Journal of Clinical Medicine· Vol 15, pp. 6005· 0 citations· 26 references
Medicine
TL;DR
It is shown that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.
Abstract
Background: The reliable integration of large language models (LLMs) into neuroimaging data extraction workflows remains unresolved. Prior benchmarking shows that exact-match accuracy underestimates LLM extraction performance, but whether inter-model consensus and variable complexity can guide automation remains unclear. We evaluated whether inter-model consensus can serve as a confidence signal for human–artificial intelligence (AI) extraction and can guide complexity-stratified workflow triage. Methods: Four frontier LLMs were queried via OpenRouter with an identical zero-shot structured prompt to extract 22 predefined variables from 91 peer-reviewed neuroimaging AI articles, yielding 2002 article–variable items per model. Variables were stratified a priori into low- (n = 7), medium- (n = 8), and high-complexity (n = 7). Performance was compared with an expert reference using exact-match and semantic-equivalence accuracy. Item-level consensus and five triage strategies characterized the efficiency–accuracy trade-off. Results: Semantic-equivalence accuracy converged to 80.5–83.4% across models despite approximately ten percentage-point exact-match differences. Unanimous 4/4 consensus occurred in 45.6% (910/1994) of items, with exact-match accuracy of 85.8%, rising to 95.3% after semantic normalization; however, 14.2% still failed to match the reference. Reliability was complexity-dependent: 96.6% for low-complexity variables, 73.2% for medium-complexity variables, and 38.1% for high-complexity variables. A hybrid strategy auto-accepting 4/4 items and routing 3/4 items to rapid verification reduced estimated review effort by approximately 59%. Conclusions: Inter-model consensus is useful, but it is incomplete and depends on variable complexity. We show that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.
A single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential.
Hyunjung Byun, Dahyoun Lee, Munyoung Jung et al.· Journal of medical systems· 1 citation· ⚡1
Accurate, context-appropriate likelihood ratios (LRs) are needed for Bayesian diagnosis, but empirical LRs are sparse because diagnostic accuracy studies are costly and context-dependent. We evaluated whether large language models (LLMs) can generate LR outputs that align with literature-reported values. We compared LR outputs generated by three OpenAI models (GPT-4o, o3, and GPT-5) with all literature-reported values curated in TheNNT.com dataset. A few-shot prompt was used to elicit numerical LR outputs, and agreement was assessed using Bland-Altman analyses to evaluate mean bias and multiplicative limits of agreement. A total of 700 reported LRs across 30 conditions were compiled, most involving signs or symptoms (59%), historical elements (19%), or test results (16%). All models demonstrated negligible mean bias. GPT-5 had the narrowest 95% limits of agreement (0.26×-3.70×) compared with o3 and GPT-4o. These findings indicate that LLM-generated LRs can be centered near literature values on average, but they are not interchangeable with published estimates for individual high-stakes decisions. Clinical use, including use for unstudied findings, would require prospective, context-specific diagnostic-accuracy validation.
Paul Chong, Shuhan He, K. Samadian et al.· Scientific Reports· 0 citations
OBJECTIVE
Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies.
METHODS
We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis.
RESULTS
Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common.
CONCLUSION
LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.
Ravi Shankar, Amaevia Lim, Xu Qian· Journal of Biomedical Inform...· 0 citations
Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy.
Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey M. Plis· 0 citations
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.
D. Kumaran, Viorica Patraucean, M. Ovsjanikov et al.· 0 citations