Skip to content
Review

Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark

Aug 2026 · 0 citations · 49 references
Computer Science

TL;DR

An expert-validated benchmark for multi-aspect, full-paper retrieval that evaluates whether retrievers can consistently recover the same paper from queries targeting its motivation, method, and experimental findings, and proposes MAPLE-Synth, a retrieval-based in-context learning pipeline that leverages OpenReview discussions and human-written query exemplars to generate realistic queries reflecting researchers' interests in different aspects of a paper.

Abstract

Scientific papers contain multiple searchable facets such as background, methods. However, many paper retrieval benchmarks merely evaluate individual query-paper relevance, while overlooking other facets of the same paper. To bridge this gap, we introduce MAPLE, an expert-validated benchmark for multi-aspect, full-paper retrieval that evaluates whether retrievers can consistently recover the same paper from queries targeting its motivation, method, and experimental findings. MAPLE contains 2,095 queries about recent ML and NLP papers, grounded in both textual and multimodal content. We further propose MAPLE-Synth, a retrieval-based in-context learning pipeline that leverages OpenReview discussions and human-written query exemplars to generate realistic queries reflecting researchers'interests in different aspects of a paper. Our expert validation shows that these queries are comparable in realism to human-written queries and highly relevant to the target papers. Experiments across lexical, scientific-domain, general-purpose text, and multimodal retrievers reveal a substantial gap between retrieving a paper from any one aspect and retrieving it from all aspects: the strongest model achieves 98.1% AnyAspect@20 but only 15.7% AllAspect@20. Experiment/result queries and table-referenced queries are particularly difficult across retrievers. Although multi-chunk aggregation improves multi-aspect paper retrieval, considerable failures persist. MAPLE provides a testbed for evaluating and developing retrievers that represent scientific papers more comprehensively.

View source

Similar papers

#natural language process... Preprint Aug 2026

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which are name ideation moves, provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scient...

Ms. S. Sharon, Tom Hope · 0 citations
Preprint Sep 2026

A Systematic Multi-Domain Evaluation of Document Retrievers

Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks...

Valentin Velev, Andreas Spitz · 0 citations
#artificial intelligence Preprint Sep 2026

Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3

Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-re...

Ryan C. Barron, Cade W. Trotter, M. Eren et al. · 0 citations
Review Open access Aug 2026

LLM aspect prediction: reviewing academic papers from different aspects with Large Language Model

Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...

Zi-Hao Hu, F. Fukumoto, Jian He et al. · 0 citations
Preprint Aug 2026

GRAFT: Graph-Distilled Generative Retrieval for Facet-Aware Scientific Literature Exploration

Scientific papers may relate by problem, method, result, or contribution, but document-level retrievers collapse these into a single similarity score without saying why they are related. Citation- and similarity-based retrieval alone also confines search to the neighbourhood of what is already known, whereas generative...

Italo Luis da Silva, Hanqi Yan, Yujing Wang et al. · 0 citations
Preprint Aug 2026

SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG

We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K pap...

Kaysarul Anas Apurba, Mahade Hasan, Rofiqul Alam Shehab et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.