SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition, reveals a dissociation between artifact quality and instructional effectiveness.
Abstract
LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.
The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.
Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen et al.· 0 citations
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological...
Kristina Šekrst, A. Kovačič· Oxford Intersections: AI in...· 0 citations
This work introduces Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation, and developed the Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks.
Zhiyi Duan, Jiangshan Guan, Qianli Xing· Proceedings of the 32nd ACM...· 0 citations
Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil...
Hao Liang, Qi-Han Lin, Mei-Yi Qiang et al.· 0 citations
Automated assessment of programming assignments remains a major challenge in higher education, particularly in large-scale courses where timely, consistent, and pedagogically meaningful feedback is difficult to provide, while Large Language Models (LLMs) have shown strong capabilities in code understanding and feedback...
Pablo Manuel Vigara Gallego, Ascensión López-Vargas, Ángel García-Beltrán et al.· Applied Sciences· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.