Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 8789-8800· 0 citations· 51 references
TL;DR
This work introduces Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation, and developed the Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks.
Abstract
The modern classroom is an inherently multimodal environment, rich with the teacher's speech, student expressions, and interactive instructional resources. Effective AI assistants should therefore perceive this complex pedagogical process, not just answer text-based questions. However, existing evaluation methods are fundamentally misaligned. Current educational benchmarks primarily focus on unimodal and single-task evaluation, while existing multimodal benchmarks focus on content evaluation by testing MLLMs as students rather than the critical process evaluation by testing them as assistants. To fill this critical gap, we introduce Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation. Inspired by educational theories, we developed our Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks. Edu-Eval is constructed from over 3,000 real-world scenarios, comprising 75,900 annotated samples. Our evaluation of eight major MLLM series reveals a critical gap between current MLLM capabilities and the demands of real-world applications. Edu-Eval is more than a dataset. It is a diagnostic tool and a clear roadmap for future research, highlighting the urgent need for the community to shift its focus from isolated content evaluation to a holistic and scenario-centric understanding of the entire classroom ecosystem.
ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.
Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al.· 0 citations
SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition, reveals a dissociation between artifact quality and instructional effectiveness.
Jing-Zhuo Wu, Jia-Jun Zhang, Liu Yi et al.· 0 citations
Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce...
Michal Štefánik, J. Nehyba, Jirina Karasova et al.· 0 citations
The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.
Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen et al.· 0 citations
Large language models (LLMs) are rapidly reshaping the technological landscape of higher education, extending from conversational assistance and personalized learning to assessment, recommendation, and institutional services. Yet their educational value cannot be inferred from generative capability alone. What matters...
Zheng-Dao Li, Jordi Conesa, Moncef Gabbouj et al.· International Journal of Edu...· 0 citations
Undergraduate thesis (UGTE) serves as a critical indicator of a student’s cumulative academic development throughout university education. Although large language models (LLMs) have advanced educational intelligence, existing approaches typically focus on holistic assessment with only a single evaluation score, ove...
Jia-Lu Zhang, Qing-Yang Sun, Qian-Yi Wang et al.· CAAI Artificial Intelligence...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.