Skip to content
Preprint

The RAT: A Unified Bayesian Model for RAG Evaluation

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

A Bayesian evaluation framework is introduced that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow, and extends to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited human judgments with cheaper automated assessments within a unified probabilistic model.

Abstract

Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow. The model distinguishes task success. Whether the user received a correct answer (from generator success) and whether the generator behaved appropriately given the retrieval outcome. We apply the framework to 27 RAG configurations across three datasets, three retrievers, and three generators, and show that the conditional decomposition reveals substantial behavioral differences between systems that appear equivalent under marginal metrics. We further analyze the annotation allocation problem, demonstrating that retrieval-success annotations are more informative than task-success annotations for estimating policy adherence, and provide an information-theoretic explanation for this asymmetry. Finally, we extend the model to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited human judgments with cheaper automated assessments within a unified probabilistic model.

View source

Similar papers

STAR: Structure-Aware Adaptive Retrieval for RAG

STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.

Yeowon Jeon, Chong-kwon Kim, Y. Choi · 0 citations
#artificial intelligence Preprint Sep 2026

Evidence Integration in Large Language Models

A distributional theory in which evidence shifts the receiver's distribution of initial answers, driven by a receiver prior weight and a candidate evidence tilt, leading to three predictions: candidates more probable to the receiver are more persuasive, receivers more readily integrate characteristic errors of their ow...

Sebastien Kawada, M. Kellis · 1 citation
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
#machine learning Preprint Sep 2026

ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs

A hybrid LLM-guided Bayesian network structure-learning framework that uses LLMs as bounded semantic guides, and shows that language-derived semantic knowledge can substantially improve scalable probabilistic structure learning when used as bounded guidance within a statistically grounded reasoning pipeline.

Jackson Hassell, Chen Shen, Estevam Hruschka · 0 citations
Conference Open access Sep 2026

Test-Time User Alignment via Bayesian Population Guidance in Subjective Tasks

In subjective tasks, different individuals can have different correct answers for the same input—the ground truth is not fixed but rather determined by each user's personal perspective. Standard foundation models suppress this individual variation by producing population-averaged predictions; conversely, few-shot in-co...

H. Ryu, J. Kang, C. Wallraven · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.