Skip to content

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

Jul 2026 · arXiv.org · Vol abs/2607.26497 · 1 citation · 28 references
Computer Science

TL;DR

Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable default, while agentic reasoning works best after ranked discovery rather than in place of it.

Abstract

Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared tier, with a margin approaching 20 points at full scale. BM25 also anchors the low-cost end of the Pareto frontier without LLM-based construction. Dense retrieval remains efficient but less accurate, whereas graph-based RAG encounters construction walls before deployment scale and its scalable variants remain below BM25 at shared tiers. Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable default, while agentic reasoning works best after ranked discovery rather than in place of it.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3

Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-re...

Ryan C. Barron, Cade W. Trotter, M. Eren et al. · 0 citations
Open access Sep 2026

Chunk Size and Retrieval Depth Optimization in a Minimal Retrieval-Augmented Generation Pipeline

Retrieval‑augmented generation (RAG) links a large language model to a set of texts. But RAG works well only when two settings are right: how big each piece of text (a "chunk") is, and how many pieces the model reads. People often choose these by habit, not by data. In this study we tested how both settings change the...

Sajjad Ahmad, Jasim Hussain, Hamza Najeeb et al. · 0 citations
Preprint Sep 2026

BoundaryMORPH: Budgeted Reranking via Active Set Selection for Diffuse Retrieval

Open-ended queries in modern Retrieval-Augmented Generation (RAG) are increasingly"diffuse,"requiring a large set of documents to be assembled into a finite LLM context window. To ensure retrieval quality, systems use fast dual-encoders and more expensive cross-encoders (CEs) to score candidates. However, the CE budget...

Eylon Caplan, Shamik Roy, S. Dasgupta et al. · 0 citations
Preprint Aug 2026

AdaWidth: Query-Adaptive Embedding Width for Dense Retrieval

AdaWidth is introduced, which adapts the number of evaluated dimensions to each query within a shared prefix representation, and derives a prefix sufficiency analysis showing that the required number of dimensions is set by the competing documents at the retrieval cutoff.

Shu-Bing Yang, Dongfang Zhao · 0 citations
#artificial intelligence Preprint Sep 2026

When Harness Beats Scale, and When Reading Beats Both

We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consiste...

Ivan Bondarenko, Nikolay O. Nikitin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.