A small, domain-specialized retrieval-augmented LLM can safely and substantially outperform far larger generalist models on a sensitive pediatric clinical domain, demonstrating that specialization plus citation-grounded retrieval, not scale alone, is a practical path to deployable clinical AI.
Background/Objectives: Large language models (LLMs) show promise for clinical decision support, yet their accuracy in interpreting specialized medical guidelines remains uncertain. Retrieval-augmented generation (RAG) may enhance performance by grounding responses in authoritative knowledge bases. This study aimed to compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions derived from the German S3 guideline for oral cavity carcinoma. Methods: We conducted a prospective, single-blind benchmark study evaluating six LLMs: one RAG-enhanced model (Custom GPT with guideline access), one consensus-based model (ConsensusGPT), and four standard models (DeepSeek-V3.2, Mistral Small 3.2, Qwen3-Next-80B, GPT-OSS-120B). Fifty clinical questions covering 17 guideline domains were presented to each model three times, yielding 900 evaluations. Three expert reviewers assessed responses using 5-point Likert scales for accuracy, comprehensiveness, and clarity, under a single-blind procedure, the effectiveness of which was tested by a pre-specified manipulation check. We then ran a paired within-model experiment in which each base model was queried with and without guideline access through a transparent, openly released retrieval pipeline, and scored every response with a condition-blind automated judge alongside deterministic retrieval metrics computed from the logs. Secondary outcomes included hallucination rates and guideline citation behavior. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Results: In a paired within-model design that held each base model fixed, adding transparent guideline retrieval improved accuracy—significantly in the three weaker open-weight models (Mistral, Qwen3, and GPT-OSS) and directionally in the already-strong DeepSeek and GPT-5 bases. Because a pre-specified blinding check found that experts could still identify retrieval-augmented answers with 98.5% accuracy, we anchored causal interpretation on measures that do not depend on the human raters, ranked by their independence: deterministic, log-derived retrieval metrics first, and then an automated, condition-blind LLM judge, whose agreement with the experts (Spearman ρ = 0.81, 95.7% within-one agreement) establishes shared calibration rather than independence from their bias. Deterministically from the retrieval logs, citation groundedness rose from 0% to 51–89% and retrieval recall@5 was 92%. On the judge, content-level hallucination fell from 42% to 4% and accuracy rose by a pooled +0.64 points (95% CI 0.47–0.80); the accuracy gain persisted after adjustment for response length (+0.48, 95% CI 0.22–0.73), which retrieval shortened rather than lengthened. The accuracy gain was large for weaker base models and small or non-significant for already-strong ones, whereas the hallucination and auditability gains were consistent across all models. The human ratings reproduced the judge’s accuracy effect (+0.61, 95% CI 0.49–0.74), and GPT-5 run through the transparent pipeline showed no significant difference from the proprietary Custom GPT (judge accuracy 4.48 vs. 4.58). Conclusions: Guideline retrieval yields a reproducible, largely base-independent improvement in the safety and auditability of LLM answers to clinical guideline questions, with accuracy gains concentrated in weaker base models. Because retrieval-augmented answers are recognizable to experts, rigorous evaluation should rely on rater-independent measures, and residual hallucination continues to require human oversight.
Andreas Vollmer, Lara Schorn, Felix Schrader et al.· Diagnostics· 0 citations
NCCN-anchored RAG outperformed both baseline GPT-5 and a literature-anchored clinical AI without direct guideline access, and external validation of guideline anchoring's operational importance provides external validation of guideline anchoring's operational importance.
D. Dukes, C. Yost, Runzhi Wang et al.· Gynecologic Oncology· 0 citations
VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Praveen Reddy, C. Mandke, Suvrankar Datta et al.· 0 citations
PURPOSE
To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners.
METHODS
Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES).
RESULTS
All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs.
CONCLUSION
LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations.
LEVEL OF EVIDENCE
III.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations
RAG-based LLM systems improved performance on anesthesiology board-style questions, but gains depended strongly on retrieval design, and reasoning-oriented models demonstrated that multistep reasoning can, in some settings, compensate for larger parameter scale.
It is suggested that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature.
P. A. Krump, Mallory N. Blasingame, T. Koonce et al.· medRxiv· 0 citations