Skip to content
Book Open access

Edu-Eval: A Large-Scale Multimodal Benchmark for MLLMs in Authentic Educational Scenarios

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 8789-8800 · 0 citations · 51 references

TL;DR

This work introduces Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation, and developed the Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks.

Abstract

The modern classroom is an inherently multimodal environment, rich with the teacher's speech, student expressions, and interactive instructional resources. Effective AI assistants should therefore perceive this complex pedagogical process, not just answer text-based questions. However, existing evaluation methods are fundamentally misaligned. Current educational benchmarks primarily focus on unimodal and single-task evaluation, while existing multimodal benchmarks focus on content evaluation by testing MLLMs as students rather than the critical process evaluation by testing them as assistants. To fill this critical gap, we introduce Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation. Inspired by educational theories, we developed our Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks. Edu-Eval is constructed from over 3,000 real-world scenarios, comprising 75,900 annotated samples. Our evaluation of eight major MLLM series reveals a critical gap between current MLLM capabilities and the demands of real-world applications. Edu-Eval is more than a dataset. It is a diagnostic tool and a clear roadmap for future research, highlighting the urgent need for the community to shift its focus from isolated content evaluation to a holistic and scenario-centric understanding of the entire classroom ecosystem.

Read PDF

Similar papers

Preprint Aug 2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.

Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al. · 0 citations
#natural language process... Preprint Sep 2026

SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition

SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition, reveals a dissociation between artifact quality and instructional effectiveness.

Jing-Zhuo Wu, Jia-Jun Zhang, Liu Yi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Edustories: A Collection of Real-world Case Studies from Classroom Practices

Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce...

Michal Štefánik, J. Nehyba, Jirina Karasova et al. · 0 citations
Preprint Aug 2026

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.

Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen et al. · 0 citations
#generative ai Open access Sep 2026

Large language models for teaching and learning in higher education: opportunities, challenges, and future directions

Large language models (LLMs) are rapidly reshaping the technological landscape of higher education, extending from conversational assistance and personalized learning to assessment, recommendation, and institutional services. Yet their educational value cannot be inferred from generative capability alone. What matters...

Zheng-Dao Li, Jordi Conesa, Moncef Gabbouj et al. · 0 citations
Open access Sep 2026

PEMUTA: Pedagogically-Enriched Multi-Granular Undergraduate Thesis Assessment

Undergraduate thesis (UGTE) serves as a critical indicator of a student’s cumulative academic development throughout university education. Although large language models (LLMs) have advanced educational intelligence, existing approaches typically focus on holistic assessment with only a single evaluation score, ove...

Jia-Lu Zhang, Qing-Yang Sun, Qian-Yi Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.