Skip to content

Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

Results show that late-interaction visual retrieval benefits from query-aware allocation rather than query-agnostic pooling alone, and identifies token top-k for interactive retrieval and the naive greedy implementation as a quality upper envelope.

Abstract

Late-interaction visual document retrievers preserve fine-grained page evidence by storing many token embeddings per page, but the resulting storage and query-time interaction costs make large-scale deployment expensive. Pooling document tokens before indexing offers a natural remedy, yet static pooling must decide which visual evidence to preserve before the query is known. We study an alternative: a heavily compressed hot-path index generates candidates, after which query-aware token budgeting operates on the original token sets of the shortlisted pages. We formulate this stage-two selection as a budgeted MaxSim coverage problem, show that a clipped version is monotone submodular, and compare coverage-only, cluster-guided, token-wise, and marginal-gain policies. On ten ViDoRe tasks with ColModernVBERT, direct static pooling reduces macro normalized discounted cumulative gain at rank five from 0.6309 without compression to 0.4738 at a thirty-two-fold pool factor. Under the same candidate-generation regime and a pool-factor-eight-equivalent reranking budget, token top-k recovers 93.93 percent of the full-token score, while greedy marginal-gain selection recovers 98.39 percent. Held-out and leave-one-dataset-out evaluations yield positive greedy improvements over token top-k on every dataset. The latency analysis reveals two useful operating points: token top-k for interactive retrieval and the naive greedy implementation as a quality upper envelope. Together, these results show that late-interaction visual retrieval benefits from query-aware allocation rather than query-agnostic pooling alone.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.

Adrien Mialland, Marc Plantevit, Julien Gallois et al. · 0 citations
#machine learning Preprint Sep 2026

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

This work revisits RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views and proposes Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agn...

T. Nguyen, Qi-Ran Hu, Ban-Ruo Liu et al. · 0 citations
Preprint Sep 2026

Test-Time Adaptation with Query-Dependent Residuals for Visual Document Retrieval

Visual document retrieval (VDR) systems depend on page embeddings computed before deployment, which makes adaptation difficult when encoder parameters or corpus re-encoding are unavailable. Rerankers provide useful relevance signals, but conventional reranking applies them only to selected queries and candidate pages....

Ze-Liang Li, Xiao-Fen Xing, K. Guo et al. · 0 citations
Preprint Aug 2026

Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals

Semantic Compression Trees is introduced, a hierarchical index in which each node stores only its semantic residual -- the information it adds beyond its parent -- and retrieval proceeds by progressive descent from the root, so that per-query cost is governed by tree depth rather than collection size.

Junaid Farooq · 0 citations
Preprint Aug 2026

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

This work proposes Trident, with two complementary components: Trident-R, a retriever-agnostic LLM reranker that converts each candidate into an LLM-readable semantic record, then performs a single adaptive-K rerank call; and Trident-S, a generation-side module that prompts the VLM under topical, entity, and structural...

Guanchen Wu, Jia-Yuan Ding, Subhabrata Mukherjee et al. · 0 citations
Preprint Sep 2026

QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage with...

Sheng-Li He, Yong-Chao Liang, Rou-Meng He et al. · 1 citation · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.