Skip to content

SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition, reveals a dissociation between artifact quality and instructional effectiveness.

Abstract

LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.

View source

Similar papers

Preprint Aug 2026

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.

Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen et al. · 0 citations
Open access Aug 2026

Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?

While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological...

Kristina Šekrst, A. Kovačič · 0 citations
Book Open access Aug 2026

Edu-Eval: A Large-Scale Multimodal Benchmark for MLLMs in Authentic Educational Scenarios

This work introduces Edu-Eval, the first large-scale benchmark focused on educational scenario-centric evaluation, and developed the Teacher-Student-Resource framework, which computationally operationalizes abstract pedagogical concepts into nine concrete evaluation tasks.

Zhiyi Duan, Jiangshan Guan, Qianli Xing · 0 citations
#natural language process... Preprint Sep 2026

OmniEdu: Open Foundation Models for Learning and Teaching

Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil...

Hao Liang, Qi-Han Lin, Mei-Yi Qiang et al. · 0 citations
Open access Sep 2026

From Black-Box Grading to Pedagogically Aligned AI Assessment: A Hybrid LLM–RAG Framework for Explainable and Scalable Automated Code Evaluation

Automated assessment of programming assignments remains a major challenge in higher education, particularly in large-scale courses where timely, consistent, and pedagogically meaningful feedback is difficult to provide, while Large Language Models (LLMs) have shown strong capabilities in code understanding and feedback...

Pablo Manuel Vigara Gallego, Ascensión López-Vargas, Ángel García-Beltrán et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.