MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs
Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and phys...