Preprint
Jul 2026
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.
Wen-Qing Zhu, Yabin Zhang, Wenjun Zeng et al.
· 0 citations