Skip to content

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

Jul 2026 · arXiv.org · Vol abs/2607.19747 · 3 citations · 92 references
Computer Science

TL;DR

Rubric4Setwise is proposed, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds, validating the effectiveness of closing the loop from evaluation to optimization.

Abstract

As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.

View source

Similar papers

Preprint Sep 2026

A Systematic Multi-Domain Evaluation of Document Retrievers

Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks...

Valentin Velev, Andreas Spitz · 0 citations
Preprint Sep 2026

Seek: Self-Evaluative Exploration for Knowledge Retrieval

seek, Self-Evaluative Exploration for Knowledge Retrieval, a training-free framework that addresses this limitation through iterative corpus interaction at test time through iterative corpus interaction at test time.

Amin Bigdeli, Radin Hamidi Rad, Negar Arabzadeh et al. · 0 citations
Preprint Aug 2026

Robustness of IR Models to Collection Growth

This study empirically evaluates the robustness of an IR model to the addition of non-relevant documents by merging two collections with negligible topic overlap and finds that MDA is more effective than MDD for retrieval, whereas MDD and MDA rerankers are equally effective.

Emmanouil Georgios Lionis, Sean MacAvaney, Debasis Ganguly · 0 citations
Preprint Sep 2026

From Ranked Documents to Reliable Contexts: An Answer-Oriented Context Construct Framework for AI Search

Traditional Web search follows a human-facing paradigm in which users inspect ranked documents and synthesize information themselves. In AI Search, retrieved documents instead serve as inputs to a generation model, shifting the retrieval objective from ranking documents by Search Satisfaction to constructing reliable c...

Yun-Fei Zhong, Yin-Qiong Cai, Lixin Su et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.