Jul 2026· International Conference on Artificial Intelligence in Education· pp. 546-561· 0 citations· 35 references
Computer Science
TL;DR
The findings suggest that Socratic guidance supports the development of students' capacity to learn with LLMs over time, highlighting its importance for LLM tutor design.
Abstract
While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education. Importantly, learning depends on how students engage with LLMs. This study examined how two types of LLM-based tutors shape students'prompting practices, learning, and subsequent LLM-use: a Socratic-Guidance (SG) tutor, which structures interaction through dialogic questioning, and a Prompt-Refinement (PR) tutor that guides the formulation of effective prompts. We conducted a two-phase study in a graduate-level mobile robotics course: 66 students used either the SG or PR tutor during a 6-week intervention, followed by 52 students using an unconstrained LLM during a 3-week course project. Results show that while the SG- and PR tutors led to similar task performance and prompting patterns during guided use, they differ in learning outcomes and later LLM-use. SG-students, relative to PR-student, achieved higher learning gains in later sessions, and were more likely to adopt understanding-driven prompting strategies, which are predictive of higher understanding, when using an unconstrained LLM. Although learners perceived the SG tutor as less efficient, the findings suggest that Socratic guidance supports the development of students'capacity to learn with LLMs over time, highlighting its importance for LLM tutor design.
Interest-based learning (IBL) is an educational approach where learners’ interests are used to contextualize learning. IBL can make instruction feel more relevant and lead to improved learning outcomes, but it is difficult for instructors to implement at scale because learner interests are highly varied. Large language models (LLMs) can support IBL through conversational AI tutors that personalize instruction to individual interests. This paper presents a prompt design approach for creating LLM tutors for IBL. We first conducted a literature review to derive pedagogy-grounded strategies for a base tutor prompt, then embedded additional IBL strategies to produce an IBL tutor prompt. We evaluated both prompts via expert review and a human-participants study with undergraduate students. Results show the IBL prompt reliably integrated learner interests, but exhibited shallow reflection, inconsistent knowledge checks, and surface-level analogies when interests were underspecified. We contribute a reusable prompt design pipeline, prompt templates, and evaluation artifacts for designing interest-based AI tutors.
Abhishek Kulkarni, S. Brown, Neha Rani et al.· International Conference on...· 0 citations
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students'subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.
Suhyeon Lee, Juneha Baek, Jaehyeong Park et al.· 0 citations
Background. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Objectives. We investigated whether AI tutors guided by SRL and CE frameworks affect conceptual understanding, perceived usability, cognitive load, student engagement, and other outcomes when compared to a pedagogically constrained baseline tutor in CS1. Methods. We conducted a preregistered, three-armed crossover study over six weeks of authentic coursework, comparing a baseline AI tutor against two SRL and CE-guided tutors designed to model Zimmerman’s cyclical model, which scaffolds planning, monitoring, and reflection, and Chi’s ICAP framework, which promotes progressively deeper forms of cognitive engagement. We assessed outcomes using post-exercise surveys, conceptual multiple-choice questions, in-platform ratings, and coded responses from the interaction logs, and analyzed the quantitative measures with mixed-effects models. Findings. On the four preregistered confirmatory measures, we found no statistically significant differences between conditions. Non-confirmatory analysis showed that students spent significantly more time on task, wrote longer messages, and produced more constructive contributions when interacting with SRL and CE tutors. Furthermore, the relationship between cognitive load and quiz performance differed significantly by agent type. Implications. Our results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment. Rather than prescribing a single tutoring strategy, future designs may benefit from giving students greater agency over the kind of help they receive, allowing them to choose between scaffolded and more direct support based on their own needs.
Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al.· Proceedings of the 2026 ACM...· 0 citations
Generative AI tutors have become a common tool for independent learning, yet their capacity to support self-regulated learning (SRL) is poorly understood. This simulation-based textual analysis of prompt design evaluates a frontier large language model (Claude Sonnet 4.6) as a tutor across 60 scripted sessions on a single topic (density), crossing three levels of SRL-informed system prompting (Minimal, Moderate, Extensive) with four learner-behavior variants (Standard, Misconception, Disengagement, Overconfidence). Tutoring transcripts were scored on a 14-dimension framework spanning SRL phases, SRL developmental stages, self-determination theory principles, and Merrill’s First Principles of Instruction, applied via an LLM judge. Adding SRL context to the system prompt raised total tutoring scores, but only at the Extensive SRL support level. Minimal and Moderate prompting produced the same performance, near 36 on a 70-point scale, and Extensive prompting raised it to 40, a statistically significant effect (partial η2 = 0.24). The learner’s behavior in the session had a larger effect than the prompt did (partial η2 = 0.37), with disengaged learners scoring lowest. The threshold pattern held under an independent judge from a different developer than the tutor model. The findings support a method for evaluating GenAI tutors empirically and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
As students increasingly use LLM tutors in computer science education, one question becomes especially important: what kind of response helps a student continue productively? Prior work has studied how students use LLMs in computer science education, but less is known about how tutoring response styles are associated with student follow-up across programming help-seeking contexts. This paper analyzes StudyChat (UMass, 2026), a public dataset of student and ChatGPT tutoring conversations from an artificial intelligence course. We transformed StudyChat into 16,851 assistant-response interactions from 203 students and 2,214 conversations. Using local LLM-assisted annotation with Gemma 4, we labeled student help-seeking situations, student state, assistant response style, and student next-turn outcome. Human validation showed 82\% agreement with the LLM-assisted labels (Cohen's $\kappa=.74$). We analyzed productive continuation and unresolved continuation across the full dataset and across help-seeking contexts. Globally, response style was significantly associated with productive continuation, $\chi^2(7)=100.39$, $p<.001$, $V=.078$, and unresolved continuation, $\chi^2(7)=125.77$, $p<.001$, $V=.087$, though effect sizes were small. Verification feedback had the highest productive-continuation rate (82.4\%), while direct answers had the lowest (62.7\%). Descriptively, response-style score ranges were smallest in low-confusion conceptual contexts (.017) and largest in high-cognitive-load contexts (.203). More detailed comparisons showed situation-dependent response patterns. For example, stepwise guidance was followed by greater confusion decrease in high-cognitive-load code requests, while direct answers were followed by more unresolved continuation in high-load debugging. These findings support context-aware evaluation and design of AI tutoring responses for programming education.
The results suggest that dialogue format should be selected according to learner characteristics in TTS dialogue-based lessons, with a significant interaction between learner characteristics and dialogue format for ARCS-based motivation.
Fumie Watanabe, Tota Suko, T. Ishida et al.· 0 citations