Skip to content
Preprint

Robustness of IR Models to Collection Growth

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

This study empirically evaluates the robustness of an IR model to the addition of non-relevant documents by merging two collections with negligible topic overlap and finds that MDA is more effective than MDD for retrieval, whereas MDD and MDA rerankers are equally effective.

Abstract

Information Retrieval (IR) systems seek to identify relevant documents within a collection. In practical applications, collections are dynamic, with documents frequently added. We argue that ideally, a retriever's effectiveness should not decrease when non-relevant documents are added to a collection. This study formalises this concept and empirically evaluates it by merging two collections with negligible topic overlap. We hypothesise that the way an IR model conditions its ranking on other documents in a collection (e.g., the IDF component in BM25 or contextual documents in listwise rerankers) plays an important role in its robustness to the addition of non-relevant documents. We broadly classify models as those that do not depend on other documents (Multi-Document-Agnostic, MDA) and those that do (Multi-Document-Dependent, MDD). Our results show that neither MDD nor MDA models are fully robust to the addition of non-relevant documents, as all models exhibit some performance degradation. Interestingly, among the models we test, MDA is more effective than MDD for retrieval, whereas MDD and MDA rerankers are equally effective.

View source

Similar papers

Preprint Sep 2026

A Systematic Multi-Domain Evaluation of Document Retrievers

Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks...

Valentin Velev, Andreas Spitz · 0 citations
Preprint Aug 2026

How retriever redundancy and diversity impact RAG effectiveness

It is shown that duplicate redundancy and LLM paraphrasing does not significantly improve answer correctness, however, providing diverse documents is highly beneficial, improving answer correctness by 17%-47%.

Jonathan J. Ross, B. Koopman, A. H. van der Vegt et al. · 2 citations
Preprint Aug 2026

Scout: Scalable Document Extraction via Data Similarity

Extracting values from large document collections powers data analysis across many domains. Frontier LLMs extract such values accurately, but processing an entire collection with one is prohibitively costly. Yet this cost is largely avoidable: real-world collections exhibit rich similarity, so for the same query over s...

Yi-Ming Lin, Chiyu Hao, Shreya Shankar et al. · 0 citations
Open access Aug 2026

Provisioning An Adaptive Model to Analyze Uncertainty and Large Language Patterns for Enhanced Document Re-Ranking

This work introduces a trust based adaptive reranking model- ATM (Adaptive Trust Model) that allocates computational resources according to file level uncertainty, instead of assigning a fixed number of reranker calls per query, which focuses computation only where ranking confidence is low.

Jenny Kalaiarasi.S · 0 citations

: Breaking the Top-k

Khaled Albishre · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.