Skip to content

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

Sep 2026 · 0 citations · 21 references
Computer Science

Abstract

Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and retrieval to the incoming query. Our approach first discovers semantic clusters over a given set of questions associated to a corpus and learns, for each cluster, a chunking strategy together with a suited metadata filtering and reranking configuration. At inference time, queries are routed to the appropriate pre-built index through nearest-centroid assignment. To further improve retrieval, we propose a supervised query router (QRe) that predicts which collections are most likely to contain relevant evidence, coupled with a Uniform Multi-source Sampler (UMS) that allocates the retrieval budget evenly across the selected sources. We evaluate our framework on large-scale, heterogeneous historical archives and show that conditioning both indexing and retrieval on the query consistently outperforms both naive baselines and strong state-of-the-art RAG systems in complex expert-domain environments.

View source

Similar papers

Preprint Sep 2026

Pre-retrieval Query Clustering for Adaptive Top-k Document Retrieval in RAG Systems

RAG systems commonly retrieve a fixed number of documents (top-k) to ground generation, but this static approach is brittle: simple queries suffer over-retrieval (adding noise and cost) while complex queries are under-retrieved, causing recall failures that cascade into incorrect answers. Motivated by the question of h...

Ye Xia, Emre Yamangil, Hai-Xun Wang · 0 citations
Conference Aug 2026

Query-Aware Multi-Relational Evidence Graph for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) systems typically apply identical retrieval logic to all user queries, overlooking that different query intents depend on fundamentally different document relationships. We propose QMEG, a query-aware retrieval framework that constructs a multi-relational evidence graph with five ac...

Jia-Run Pan, Yu-Ling Fan, Li Ma et al. · 0 citations
#natural language process... Preprint Sep 2026

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

CoGR is introduced, a retrieval framework that instead trains LLMs to directly construct retrieval representations on both query and item sides, and shows stable co-evolution and increasingly aligned query--item keyword spaces over training.

Runpeng Dai, Kai-Li Huang, Changsung Kang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query complexity and information needs, leading to inefficient allocation of computational resources. While retrieval and generation adaptivity have been studied...

Neeraj Anand, Payel Santra, Partha Basuchowdhuri et al. · 0 citations
Preprint Aug 2026

D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding

The core of D2-ScaleAgent is a Verifier agent-driven dynamic routing loop based on the intrinsic difficulty of the query, centered around a continuously updated evidence bank that serves as the agent's dynamic working memory.

Hao Zhang, Longrong Yang, Lun-Hao Duan et al. · 0 citations
Jul 2026

GAS: A Lightweight Framework for Filtered Search over Wide-Table Vectors

Wide-table vectors, where each embedding is linked with numerous structured attributes, are prevalent in applications such as autonomous driving and multimodal data processing for large-model training. Efficiently retrieving semantically similar vectors under attribute filters is crucial for these tasks, a problem addr...

Zi-Yuan He, Yu-Xiang Wang, Yu Sun et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.