Skip to content
Preprint

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.

Abstract

Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.

View source

Similar papers

#artificial intelligence Review Oct 2026

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review...

Zhang Jiang, Zina Ibrahim, J. Teo · 0 citations
#artificial intelligence Review Aug 2026

Medical Causal Hypothesis Verification with Large Language Models

It is shown that while LLMs exhibit strong recall, they often perform poorly at providing valid scientific articles and evidence for support and at rejecting unsupported hypotheses, highlighting the need for rigorous evaluation before using LLMs for search and retrieval in healthcare settings.

Safiyyah Ahmed, Abrar Ansari, M. Islam et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support

Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation Stability Index (ES...

B. Bolla, Vishnu Surya Reddy Nandi · 0 citations
Preprint Aug 2026

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

A two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains, and shows that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidenc...

Uma S. Ranjan, Kunal Tilaganji, Aditya Koul et al. · 0 citations
Review Open access Sep 2026

Citation reliability of frontier large language models in medical writing and its automated verification

Large language models are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them, so LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.

R. Shin, J.-M. Lee, J. Park et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.