Skip to content
Preprint

A bottom-up taxonomy of student discourse with a Socratic AI physics tutor

Aug 2026 · 0 citations · 2 references
Physics

TL;DR

A description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor.

Abstract

Large language model (LLM) tutors are being deployed in introductory physics courses at a scale that produces transcript corpora far larger than traditional qualitative coding can absorb. A central question for physics education research (PER) is empirical and prior to any claim about effectiveness: what do students actually say to these tutors? We address this question for one Socratic AI tutor deployed in an introductory calculus-based mechanics course by building a bottom-up taxonomy of student discourse. Each student turn is assigned an emergent free-text label by an LLM coder using the surrounding conversational context; near-paraphrase labels are then consolidated into a smaller set of discourse categories using a similarity-based grouping procedure. The procedure is validated against a stratified human-coded sample. The resulting taxonomy of 357 categories is strikingly concentrated: the top 25 categories cover roughly half of all student turns, and two thematic bands: equation-handling and meta-procedural requests together dominate the head of the distribution. The substantive contribution is the taxonomy itself: a description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor

View source

Similar papers

Book Open access Jul 2026

A Good Rubber Duck Does Not Quack: Designing Socratic Scaffolding in AI Tutors

Socratic AI, a VS Code-integrated tutor that addresses this through pedagogically-grounded Socratic dialogue constrained to withhold direct solutions is presented, and evidence that stateful tracking enables adaptive Socratic dialogue that scaffolds productive struggle rather than short-circuiting learning is contributed.

Ayush Thonge, Aalok Thakkar · 0 citations
Open access Jul 2026

The Socratic Trap: Benchmarking the Capacity of Large Language Models to Generate Strategic Misconceptions in Computer Science Education

SocraticTrap-CS is introduced, a publicly available benchmark that probes the capacity of open-weight LLMs to generate strategic misconceptions on demand and reframes the evaluation of educational LLMs around pedagogical trustworthiness rather than factual correctness alone.

Marijela Miličević, Mia Rovis, Ratomir Karlović et al. · 0 citations
Open access Jul 2026

Prompting for Independent Learning: An Evaluation of Tutoring Behaviors in GenAI

Generative AI tutors have become a common tool for independent learning, yet their capacity to support self-regulated learning (SRL) is poorly understood. This simulation-based textual analysis of prompt design evaluates a frontier large language model (Claude Sonnet 4.6) as a tutor across 60 scripted sessions on a single topic (density), crossing three levels of SRL-informed system prompting (Minimal, Moderate, Extensive) with four learner-behavior variants (Standard, Misconception, Disengagement, Overconfidence). Tutoring transcripts were scored on a 14-dimension framework spanning SRL phases, SRL developmental stages, self-determination theory principles, and Merrill’s First Principles of Instruction, applied via an LLM judge. Adding SRL context to the system prompt raised total tutoring scores, but only at the Extensive SRL support level. Minimal and Moderate prompting produced the same performance, near 36 on a 70-point scale, and Extensive prompting raised it to 40, a statistically significant effect (partial η2 = 0.24). The learner’s behavior in the session had a larger effect than the prompt did (partial η2 = 0.37), with disengaged learners scoring lowest. The threshold pattern held under an independent judge from a different developer than the tutor model. The findings support a method for evaluating GenAI tutors empirically and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena · 0 citations
Review Jul 2026

P-A.I.R.: A Structured AI-Replication Framework for Active Learning in Introductory Physics

The growing use of generative AI tools among students raises an important pedagogical question: how can AI be structured to promote active learning rather than passive answer-seeking? This study introduces the Physics AI-Replication (P-A.I.R.) framework, in which students identify challenging problems, use AI to explain underlying concepts and solutions, generate similar problems, and practice independently before reviewing answers. Survey data from 39 undergraduate students in algebra-based physics indicate that nearly all participants reported improved conceptual understanding following engagement with P-A.I.R. across the semester, and self-confidence scores were consistently above the scale midpoint (mean = 7.03; range 5-9). A strong association between conceptual clarity and replication helpfulness (Spearman r = 0.582, p<0.001) supports the framework's design. Qualitative findings highlight conceptual clarification and structured problem-solving. These results suggest that P-A.I.R. offers a practical approach for integrating AI into physics learning.

Bilas Paul · 0 citations
Review Open access 2026

BEYOND THE “SAGE ON THE STAGE”: GENERATIVE AI AND THE CHALLENGING ROLES IN ESP TEACHING

The massive spread of Generative Artificial Intelligence (Gen AI) into language classrooms has reopened a currently relevant question related with the human teacher’s role, and in some quarters revived the worry that the instructor who once stood as the “Sage on the Stage” has become obviously less important. This narrative literature review asks what actually happens to the roles and professional identities of language teachers as Gen AI has massively intervened their work, with particular attention to English for Specific Purposes (ESP). Following a protocol-guided search of academic databases, thirteen peer-reviewed studies published between 2022 and 2026 were synthesized, read alongside a smaller body of abundant literature on general language teaching. Taken together, the studies point away from displacement and toward reinvention. Teachers, ESP practitioners in particular, are taking on the work of prompt design, drawing on what one study terms AI-pedagogical knowledge to translate tacit expertise into instructions a model can follow. Since Gen AI can produce credible but incorrect content in specific fields such as engineering, medicine, law, etc., ESP teachers also find themselves play roles as domain validators and as guides to critical AI literacy. What emerges is less an authoritative source of knowledge than a facilitator who decides when to lean on automation and when to rely on judgment, context, and rapport that a model cannot supply.

Gregorius Punto Aji, Angelina Kusuma Jelita Mawarni · 0 citations
Preprint Aug 2026

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.

Unggi Lee, Sookbun Lee, Yeil Jeong et al. · 0 citations