Skip to content

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

Jul 2026 · arXiv.org · Vol abs/2607.24010 · 3 citations · ⚡ 1 influential · 33 references
Computer Science

TL;DR

Budget-aware evaluation for Active RAG evaluation is studied by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer.

Abstract

Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations often leave the operating point underspecified: two systems may both claim a 50% evidence-usage budget while realizing different held-out usage rates, so higher accuracy can reflect a looser budget rather than a better retrieval policy. We study budget-aware evaluation for Active RAG by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer. This view separates three questions that single-point evaluations conflate: whether trigger scores rank useful retrieval decisions, whether thresholds calibrated on past data meet future budgets, and how trigger-side computation changes deployment cost. We operationalize these questions with exact top-k utility frontiers, deployable threshold frontiers, conservative budget frontiers, harm audits, and cost decompositions. Across knowledge-intensive multi-hop QA datasets and open instruction models, retrieval harm is non-negligible, router rankings change across datasets and budgets, nominal thresholds can miss target usage, and simple uncertainty or retrieval-score baselines often rival learned utility routers. Budget-aware Active RAG evaluations should therefore report frontiers, realized usage, threshold-transfer error, harm rates, and cost decompositions alongside accuracy.

View source

Similar papers

STAR: Structure-Aware Adaptive Retrieval for RAG

STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.

Yeowon Jeon, Chong-kwon Kim, Y. Choi · 0 citations
#artificial intelligence Preprint Sep 2026

BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence

Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow....

Hong-Ji Pu · 0 citations
#artificial intelligence Preprint Sep 2026

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We i...

Priyanath Maji, S. Chowdhury · 0 citations
Preprint Oct 2026

Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers...

Jing-Jie Ning, Xue-Qi Li, Yi-Bo Kong · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

Retrieval-Grounded Voting (RGV), which scores each rollout by the lexical overlap between its final answer and the documents it retrieved, consistently outperforms confidence-based voting and identifies the underlying failure reason as copy inflation.

Hyunho Kook, Junhyuk So, Tianyu Fu et al. · 0 citations
Preprint Aug 2026

When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1

Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem rather than an independent state-classification task. We adapt S2G-RAG's structured sufficie...

Wei Luo · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.