Skip to content
Preprint

SAGE: Semantic Attribute Graphs for Multi-Entity Visual Retrieval

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

SAGE, a training-free framework that parses semantic entities from dense document images, represents them as hierarchical graph nodes with multi-vector embeddings, and retrieves query-relevant evidence through iterative entity-level subgraph matching, is proposed.

Abstract

Dense document images often contain many fine-grained visual and textual entities whose relevance depends on a user query. Standard vision-language retrievers encode cropped regions with a single vector, which can mix distinct entity signals and obscure the evidence needed for fine-grained retrieval. We call this failure mode Semantic Dilution and quantitatively show that it degrades entity-level retrieval as a function of entity density. To mitigate it, we propose SAGE, a training-free framework that parses semantic entities from dense document images, represents them as hierarchical graph nodes with multi-vector embeddings, and retrieves query-relevant evidence through iterative entity-level subgraph matching. We also introduce DEAR, a dataset of 1,055 query--image pairs sourced from product detail pages, where each query requires retrieving and comparing multiple fine-grained entities from visually dense inputs across four question types of increasing complexity. Experiments show that SAGE substantially reduces semantic dilution and outperforms patch-level and OCR-based retrieval baselines on DEAR, achieving a Recall@3 of 0.849 and a generation score of 2.746 on multi-entity visual comparison queries. Our code is available at https://github.com/All4Nothing/SAGE.

View source

Similar papers

Preprint Aug 2026

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Hangrui Xu, Zheng-Xian Wu, Yu Yu et al. · 0 citations

CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval in Image Archives

The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.

Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al. · 0 citations
Preprint Aug 2026

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

M UL V EC is proposed, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear.

Zihao Zhang, Da-Yan Wu, Xin-Ze Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.

Adrien Mialland, Marc Plantevit, Julien Gallois et al. · 0 citations
Preprint Aug 2026

Query-Driven Multimodal Information Extraction from Long Documents

This work proposes query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxes, and designed a two-level taxonomy that operates at the query and instance levels.

Yi-Zhou Gao, Ding Xia, Xi Yang · 0 citations
Preprint Aug 2026

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

This work proposes Trident, with two complementary components: Trident-R, a retriever-agnostic LLM reranker that converts each candidate into an LLM-readable semantic record, then performs a single adaptive-K rerank call; and Trident-S, a generation-side module that prompts the VLM under topical, entity, and structural...

Guanchen Wu, Jia-Yuan Ding, Subhabrata Mukherjee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.