Skip to content
Book Open access

MMID: Multi-turn Multimodal Interactive Dialogue Benchmark

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 10 references

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal understanding and reasoning by integrating linguistic and visual information, while benchmarks facilitate iterative model improvement by evaluating their performance and analyzing their limitations. However, existing dialogue-based multimodal benchmarks do not fully reflect the characteristics of real-world interactions, as they often construct a single, lengthy user utterance to provide all requirements or treat visual information as static even in multi-turn conversations. To address these limitations, we propose the Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark, where user requirements are incrementally conveyed across turns and images are interleaved with text throughout the conversation to enable dynamic multimodal interaction. With this design, MMID enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs. Furthermore, while most tasks adopt a multiple-choice question format, each incorrect option is mapped to fine-grained error types, enabling an analysis of model strengths and weaknesses beyond coarse-grained performance comparison. MMID reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues. Our benchmarks and detailed descriptions are available at https://github.com/KUNLP/MMID.

Read PDF