Skip to content
Preprint

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

MULVEC is proposed, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear.

Abstract

Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existing methods typically describe the target as a whole and match this description with a global image representation. This global matching can mix different semantic cues and lose fine- grained details. We propose MULVEC, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear. Frozen encoders map the query to one target description vector and role-specific probe vectors, while each candidate is represented by one global visual vector and a bank of local visual vectors. The retrieval roles then use this shared evidence for their respective purposes, and a fixed weighted sum of their scores ranks the entire gallery in a single retrieval pass. Across CIRCO, CIRR, and FashionIQ and three backbone scales, MULVEC improves CIRCO mAP@5 by up to 23.0% over the strongest compared method and gives the best CIRR and FashionIQ results in our comparison.

View source

Similar papers

Preprint Aug 2026

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

M UL V EC is proposed, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear.

Zihao Zhang, Da-Yan Wu, Xin-Ze Liu et al. · 0 citations
Conference Aug 2026

StaG-CoTMR: Paraphrase-Consistent Zero-Shot Composed Image Retrieval via Finite Edit Graphs and Consensus-Guided Ranking

Zero-shot composed image retrieval (ZS-CIR) methods often use large vision-language models (LVLMs) to represent an image-text query comprising a reference image and a modification instruction. However, semantically equivalent instructions can alter the generated representations and final rankings. An individual retriev...

Bo-Wen Fu, Na Liu, Yue-Ming Shu et al. · 0 citations
Preprint Sep 2026

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constr...

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar et al. · 0 citations
Sep 2026

BPE-Level Visual-Textual Alignment for Multi-Scene Text Retrieval

Scene Text Retrieval (STR) aims to search images containing a given textual query within large-scale image collections. However, existing approaches are fundamentally constrained in two ways: 1) they are evaluated on narrow benchmarks that focus primarily on natural scenes; and 2) they rely on either error-prone multi-...

Tong-Kun Guan, Yu-Tong Cai, Hao-Cheng Wang et al. · 0 citations
Preprint Sep 2026

MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders

Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual encoder's global ima...

Abdulmalik Alquwayfili, Faisal Almeshal, Jumanah Almajnouni et al. · 0 citations
Preprint Sep 2026

Preserve-and-Compose Training for Composed Image Retrieval

Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captio...

Sehyun Kwon · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.