A description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor.
Abstract
Large language model (LLM) tutors are being deployed in introductory physics courses at a scale that produces transcript corpora far larger than traditional qualitative coding can absorb. A central question for physics education research (PER) is empirical and prior to any claim about effectiveness: what do students actually say to these tutors? We address this question for one Socratic AI tutor deployed in an introductory calculus-based mechanics course by building a bottom-up taxonomy of student discourse. Each student turn is assigned an emergent free-text label by an LLM coder using the surrounding conversational context; near-paraphrase labels are then consolidated into a smaller set of discourse categories using a similarity-based grouping procedure. The procedure is validated against a stratified human-coded sample. The resulting taxonomy of 357 categories is strikingly concentrated: the top 25 categories cover roughly half of all student turns, and two thematic bands: equation-handling and meta-procedural requests together dominate the head of the distribution. The substantive contribution is the taxonomy itself: a description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor
Socratic AI, a VS Code-integrated tutor that addresses this through pedagogically-grounded Socratic dialogue constrained to withhold direct solutions is presented, and evidence that stateful tracking enables adaptive Socratic dialogue that scaffolds productive struggle rather than short-circuiting learning is contributed.
Ayush Thonge, Aalok Thakkar· Annual Conference on Innovat...· 0 citations
SocraticTrap-CS is introduced, a publicly available benchmark that probes the capacity of open-weight LLMs to generate strategic misconceptions on demand and reframes the evaluation of educational LLMs around pedagogical trustworthiness rather than factual correctness alone.
Marijela Miličević, Mia Rovis, Ratomir Karlović et al.· Information· 0 citations
Generative AI tutors have become a common tool for independent learning, yet their capacity to support self-regulated learning (SRL) is poorly understood. This simulation-based textual analysis of prompt design evaluates a frontier large language model (Claude Sonnet 4.6) as a tutor across 60 scripted sessions on a single topic (density), crossing three levels of SRL-informed system prompting (Minimal, Moderate, Extensive) with four learner-behavior variants (Standard, Misconception, Disengagement, Overconfidence). Tutoring transcripts were scored on a 14-dimension framework spanning SRL phases, SRL developmental stages, self-determination theory principles, and Merrill’s First Principles of Instruction, applied via an LLM judge. Adding SRL context to the system prompt raised total tutoring scores, but only at the Extensive SRL support level. Minimal and Moderate prompting produced the same performance, near 36 on a 70-point scale, and Extensive prompting raised it to 40, a statistically significant effect (partial η2 = 0.24). The learner’s behavior in the session had a larger effect than the prompt did (partial η2 = 0.37), with disengaged learners scoring lowest. The threshold pattern held under an independent judge from a different developer than the tutor model. The findings support a method for evaluating GenAI tutors empirically and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
The growing use of generative AI tools among students raises an important pedagogical question: how can AI be structured to promote active learning rather than passive answer-seeking? This study introduces the Physics AI-Replication (P-A.I.R.) framework, in which students identify challenging problems, use AI to explain underlying concepts and solutions, generate similar problems, and practice independently before reviewing answers. Survey data from 39 undergraduate students in algebra-based physics indicate that nearly all participants reported improved conceptual understanding following engagement with P-A.I.R. across the semester, and self-confidence scores were consistently above the scale midpoint (mean = 7.03; range 5-9). A strong association between conceptual clarity and replication helpfulness (Spearman r = 0.582, p<0.001) supports the framework's design. Qualitative findings highlight conceptual clarification and structured problem-solving. These results suggest that P-A.I.R. offers a practical approach for integrating AI into physics learning.
The massive spread of Generative Artificial Intelligence (Gen AI) into language classrooms has reopened a currently relevant question related with the human teacher’s role, and in some quarters revived the worry that the instructor who once stood as the “Sage on the Stage” has become obviously less important. This narrative literature review asks what actually happens to the roles and professional identities of language teachers as Gen AI has massively intervened their work, with particular attention to English for Specific Purposes (ESP). Following a protocol-guided search of academic databases, thirteen peer-reviewed studies published between 2022 and 2026 were synthesized, read alongside a smaller body of abundant literature on general language teaching. Taken together, the studies point away from displacement and toward reinvention. Teachers, ESP practitioners in particular, are taking on the work of prompt design, drawing on what one study terms AI-pedagogical knowledge to translate tacit expertise into instructions a model can follow. Since Gen AI can produce credible but incorrect content in specific fields such as engineering, medicine, law, etc., ESP teachers also find themselves play roles as domain validators and as guides to critical AI literacy. What emerges is less an authoritative source of knowledge than a facilitator who decides when to lean on automation and when to rely on judgment, context, and rapport that a model cannot supply.
Gregorius Punto Aji, Angelina Kusuma Jelita Mawarni· International Journal of Edu...· 0 citations
EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
Unggi Lee, Sookbun Lee, Yeil Jeong et al.· 0 citations