Skip to content
Preprint

EXCISE: Query-Side Exclusion for Late-Interaction Retrieval

Aug 2026 · 0 citations · 45 references
Computer Science

TL;DR

Across six collections and three backbones, EXCISE is the strongest system in all eighteen backbone-collection cells against that backbone's own frozen and fine-tuned baselines, and outperforms every fine-tuned cross-encoder, each of which loses no-harm nDCG@10.

Abstract

Late-interaction retrievers handle exclusion queries poorly. When a user asks for X but not Z, the additive MaxSim score promotes documents covering Z, a problem we call exclusion inversion. We show that no readout of the frozen vectors recovers the constraint, because the difficulty lies in identifying the excluded topic, which depends on the query alone. EXCISE operates at query time and corrects the inversion while leaving the index frozen. Two query-side modules totalling 1.5M parameters identify the topic and re-embed a 100-document shortlist, and a parameter-free rule demotes candidates matching that topic. Across six collections and three backbones, EXCISE is the strongest system in all eighteen backbone-collection cells against that backbone's own frozen and fine-tuned baselines. It raises exclusion success@10 on ExcluIR from 0.058 to 0.691 and raises Boolean NOT accuracy from 0.25-0.29 to 0.90-0.92. Pooled over 1,860 queries, it outperforms every fine-tuned cross-encoder, each of which loses no-harm nDCG@10, whereas EXCISE matches its frozen baseline on its strongest backbone. We release X-BENCH, a tiered benchmark of explicit, implicit, and compound exclusions with no-harm and Boolean controls.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3

Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-re...

Ryan C. Barron, C. Trotter, M. Eren et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval

Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond...

Truong Son Nguyen, D. Blackley, Ni Trieu et al. · 0 citations
Preprint Sep 2026

Top-K Is Not a Budget for Hybrid Retrieval

Modern hybrid retrieval for RAG typically fuses the Top-$L$ results from dense and sparse retrievers, but a fixed truncation depth may not transfer across changing queries and corpora. Exact fusion removes the dependence on a fixed depth, yet completing a specified Top-$K$ still incurs variable access costs. We present...

Chunran Zhang · 0 citations
#artificial intelligence Preprint Aug 2026

Efficient GPU Retrieval for Semantic Search

A policy-aligned retrieval framework that improves offline relevance over a matched-capacity baseline, with gains broadly distributed across facet combinations, and serves this framework with a two-stage GPU architecture.

Dhritiman Das, Chujie Zheng, Ronak Kaoshik et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.