Skip to content
Review Open access

LitBench: A Benchmark for Retrieval-Grounded Multi-Paper Evidence Synthesis in Epilepsy

Sep 2026 · medRxiv · 0 citations · 25 references
Medicine

TL;DR

No system was reliable across every axis tested: accuracy fell as competing papers were added, interpretation-requiring facts were harder than stated ones, two-paper synthesis was rarely achieved, and only some systems recognized absent evidence.

Abstract

Background. Clinicians increasingly use AI systems to search the medical literature, but current benchmarks do not jointly test whether a response identifies the originating paper and its supporting passage. LitBench evaluates four dimensions: stated versus interpretation-requiring facts, the number and composition of competing papers, one- versus two-paper evidence, and refusal when evidence is absent. Methods. LitBench comprises 1,980 open-access papers (1,000 epilepsy papers holding the answers, 980 decoys from unrelated fields) and 9,472 human-reviewed facts, giving 2,188 single-paper questions and 170 requiring a fact from each of two papers. Difficulty rose by burying the answer among up to 2,000 distractors, then again over live PubMed Central. The same questions were then asked with the answering paper removed, so that refusal was the correct response. Four systems were tested: Gemma-4B, Gemma-12B, Sonnet-5, and DeepSeek-V4-Flash (refusal conditions only). Three model judges scored each answer by majority; two-paper questions counted only when both facts were found. Results. Across fixed-corpus single-paper conditions, accuracy ranged from 55.9% to 71.9% for Gemma-4B, 61.2% to 75.5% for Gemma-12B, and 90.0% to 91.3% for Sonnet-5. Similar epilepsy papers were harder than mixed candidate sets for both Gemma configurations. Two-paper accuracy was 5.7%, 17.1%, and 37.8%, respectively. With evidence absent, Gemma-12B refusal fell from 88.3% to 39.9% as candidate sets grew; Gemma-4B almost never refused. Sonnet-5 refused on 94.9% of sampled single-paper and all sampled two-paper cells. DeepSeek-V4-Flash also refused frequently, but on 80.3% of matched answer-present controls. Conclusions. No system was reliable across every axis tested: accuracy fell as competing papers were added, interpretation-requiring facts were harder than stated ones, two-paper synthesis was rarely achieved, and only some systems recognized absent evidence. What counts as "successful" literature search carries considerable nuance; LitBench can evaluate proposed tools at a fine-grained level.

Read PDF

Similar papers

Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves gro...

P. Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations
Preprint Sep 2026

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Across six medical multimodal MCQ datasets, this work separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key, showing that medical image-reasoning claims require route-level evidence.

Ben Wang, Yi-Fan Zhang, Jia-Qing Yu et al. · 0 citations
Review Open access Sep 2026

Citation reliability of frontier large language models in medical writing and its automated verification

Large language models are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them, so LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.

R. Shin, J.-M. Lee, J. Park et al. · 0 citations
Review Open access Sep 2026

Large Language Model versus Clinician Written Summaries of Research Papers.

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.

Richard Guthmann, Robert Martin, Erin Lee et al. · 0 citations
#artificial intelligence Review Oct 2026

Answering clinicians'questions over trial evidence tables with verifiable, feedback-driven language models

Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern attributes that the table does not record, such as a drug's target class or a harmonised endpoint. Here we introduce FD-SCoPE, a language-model framewor...

M. Choudhury, Suparno Roy Chowdhury, S. Sahoo et al. · 0 citations
#natural language process... Preprint Sep 2026

SentryLine: Evidence-Grounded Question Answering over Evolving Documents in Oncology Care

SENTRYLINE is presented, a living guideline-aware clinical question answering system that retrieves guideline passages through a vectorless hierarchical RAG pipeline and returns a role-specific answer with inline citations, factual and temporal verification reports, and drift detection notes that surface when a guideli...

Tampu Ravi Kumar, Gaurav Najpande, Muhammad Ali Khan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.