Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 670-675· 0 citations· 15 references
Abstract
Conversational Image Retrieval (CIR) refines image search through multi-turn interaction, where the Questioner plays a central role in eliciting information about the user’s target. However, existing CIR systems are commonly evaluated in end-to-end settings, making it difficult to determine whether performance gains originate from questioning strategies, retrieval backbones, answering behavior, or interaction protocols. We introduce a controlled benchmark for studying questioning strategies in CIR. The proposed benchmark isolates the Questioner module by keeping the retriever, answerer, dialogue budget, decoding setting, and evaluation pipeline fixed across all methods. Under this unified protocol, we compare representative strategies including Blind, Text-based, Text-based Recon, Top-K Guided, and Hybrid. Experimental analysis reveals a clear stage-dependent trade-off between efficiency and discriminative capability. Text-based strategies remain computationally lightweight but become increasingly vulnerable to contextual drift over extended interaction, whereas Top-K Guided improves retrieval refinement at the cost of higher inference latency. Overall, the proposed benchmark provides a reproducible framework for analyzing conversational questioning behavior in CIR systems.
The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.
S. Han, Thang Vu, Junyeong Kim· IEEE Access· 0 citations
In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixe...
Tam Le Thi Thanh, Van Tran Hoang, Hong-Hanh Nguyen-Le et al.· 0 citations
Multi-turn interactions with LLMs are becoming increasingly common in information-seeking scenarios. However, user queries are often ambiguous and context-dependent, making them ill-suited for direct use as retriever queries. Conversational query reformulation (CQR) addresses this issue by rewriting the current utteran...
Yunah Jang, Kang-il Lee, Joongbo Shin et al.· 0 citations
Interactive text-to-image retrieval seeks to identify a target image through multi-turn dialogue, where users progressively refine ambiguous search intent with additional visual constraints. Existing approaches commonly encode the full dialogue as a single concatenated sequence, but this strategy is vulnerable to token...
Khiem Le, Bao Tran, Duy-Dinh Le et al.· International Conference on...· 0 citations
The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-language response. This paper presents the challenge task, dataset,...
Seungheon Doh, Sergio Oramas, B. Sguerra et al.· Proceedings of the Workshop...· 0 citations
Recent Multimodal Large Language Models (MLLMs) support retrieval-augmented generation (RAG) for document question answering (QA), yet it remains unclear how effectively they use the provided evidence during answer generation. In this work, we conduct a controlled empirical study of 7 popular MLLMs on long multimodal m...
Debolena Basak, Digbalay Bose, Koustava Goswami et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.