Skip to content
Preprint

A Systematic Multi-Domain Evaluation of Document Retrievers

Sep 2026 · 0 citations · 95 references
Computer Science

Abstract

Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers'failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.