Skip to content
Conference

Questioning Matters: A Controlled Study of Question Generation in Conversational Image Retrieval

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 670-675 · 0 citations · 15 references

Abstract

Conversational Image Retrieval (CIR) refines image search through multi-turn interaction, where the Questioner plays a central role in eliciting information about the user’s target. However, existing CIR systems are commonly evaluated in end-to-end settings, making it difficult to determine whether performance gains originate from questioning strategies, retrieval backbones, answering behavior, or interaction protocols. We introduce a controlled benchmark for studying questioning strategies in CIR. The proposed benchmark isolates the Questioner module by keeping the retriever, answerer, dialogue budget, decoding setting, and evaluation pipeline fixed across all methods. Under this unified protocol, we compare representative strategies including Blind, Text-based, Text-based Recon, Top-K Guided, and Hybrid. Experimental analysis reveals a clear stage-dependent trade-off between efficiency and discriminative capability. Text-based strategies remain computationally lightweight but become increasingly vulnerable to contextual drift over extended interaction, whereas Top-K Guided improves retrieval refinement at the cost of higher inference latency. Overall, the proposed benchmark provides a reproducible framework for analyzing conversational questioning behavior in CIR systems.

View source

Similar papers

Open access 2026

Speaking ITM’s Language: Query Reformulation and Temporal-MMR for Long-Video QA

The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.

S. Han, Thang Vu, Junyeong Kim · 0 citations
#artificial intelligence Preprint Sep 2026

When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models

In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixe...

Tam Le Thi Thanh, Van Tran Hoang, Hong-Hanh Nguyen-Le et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CompCQR: Compositional Query Generation for Training-Free Conversational Search

Multi-turn interactions with LLMs are becoming increasingly common in information-seeking scenarios. However, user queries are often ambiguous and context-dependent, making them ill-suited for direct use as retriever queries. Conversational query reformulation (CQR) addresses this issue by rewriting the current utteran...

Yunah Jang, Kang-il Lee, Joongbo Shin et al. · 0 citations
Conference Aug 2026

MGTE: A Modular Multi-Granularity Text Ensemble for Interactive Image Retrieval

Interactive text-to-image retrieval seeks to identify a target image through multi-turn dialogue, where users progressively refine ambiguous search intent with additional visual constraints. Existing approaches commonly encode the full dialogue as a single concatenated sequence, but this strategy is vulnerable to token...

Khiem Le, Bao Tran, Duy-Dinh Le et al. · 0 citations
#large language models Book Open access Sep 2026

Overview and Analysis of the RecSys Challenge 2026: Conversational Music Recommendation

The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-language response. This paper presents the challenge task, dataset,...

Seungheon Doh, Sergio Oramas, B. Sguerra et al. · 0 citations
Book Open access Aug 2026

Diagnosing Evidence Utilization in Multimodal Document Question Answering

Recent Multimodal Large Language Models (MLLMs) support retrieval-augmented generation (RAG) for document question answering (QA), yet it remains unclear how effectively they use the provided evidence during answer generation. In this work, we conduct a controlled empirical study of 7 popular MLLMs on long multimodal m...

Debolena Basak, Digbalay Bose, Koustava Goswami et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.