A GenAI-Based Adaptive Tutoring ana Intelligent Assessment Framework for Personalized Learning
Dhyan GowdaM. G. ArunaPriya Ramesh PrasadPunith Kumar
S. N.
Aug 2026· International Journal of Sciences and Innovation Engineering· 0 citations
TL;DR
EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
Abstract
Modern academic institutions face a fundamental gap: instruction is designed for the average student, leaving individuals with specific weaknesses without any targeted support mechanism. Standardised course pipelines and identical assessments for all students have consistently failed to close this instructional gap. Advances across Artificial Intelligence (AI), Natural Language Processing (NLP), and Generative AI now make it feasible to construct learning environments that actively evolves as each student progresses [1].
This paper introduces EduMind, a unified tutoring and assessment platform designed around a dual-track evaluation model. Closed- form questions are evaluated using fixed-logic scoring for consistent results, while open-ended answers are assessed by computing meaning-level correspondence with expert reference responses. The combined output enables fine-grained identification of both proficient and deficient knowledge areas at the topic level [2][3].
EduMind's embedded Generative Al module transforms evaluation data directly into targeted instructional content, addressing identified gaps with structured explanations and study material packaged into downloadable PDF reports. The system demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
: Traditional educational assessment systems often prioritize grading over learning, falling short in automating the evaluation of complex coding and subjective responses. This introduces inconsistencies and slows the feedback cycle. This paper introduces InsightEval, an AI-powered system designed to transform static assessments into dynamic tools for continuous learning. It automates the evaluation of multiple-choice, coding, and subjective questions by integrating technologies like Judge0 for code execution and Large Language Models for natural language grading. The system’s core innovation is a 'Feedback -to-Improvement Loop,' which analyzes incorrect answers using NLP to identify conceptual gaps and then generates personalized micro-quizzes for targeted reinforcement. It also visualizes topic interconnections through Concept Linkage Maps, helping learners and teachers track conceptual mastery. InsightEval delivers an interactive, feedback-driven experience that promotes deeper understanding and continuous academic growth.
S. A, B. Reddy, Eric Varghese et al.· Proceedings of the 1st Inter...· 0 citations
The integration of Artificial Intelligence (AI) into higher education offers scalable support for students but raises concerns regarding over-reliance, reduced effort, and diminished deep learning. This study introduces Michael, a syllabus-aware AI teaching assistant designed to scaffold reasoning through structured, hint-first dialogue aligned with course progression, rather than providing direct solutions. The system was deployed in an undergraduate Structured Query Language (SQL) course across three consecutive semesters and evaluated using a mixed-methods design combining interaction logs, pre–post questionnaires (N = 170), and classroom observations. Results indicate high perceived ease of use (M = 4.43) and a moderate but statistically significant increase in trust following exposure (from M = 3.29 to M = 3.58), while AI self-efficacy showed only minor changes. Usage patterns revealed a bifurcated structure, with students engaging in both short troubleshooting interactions and extended tutoring dialogues. Qualitative findings highlight adoption waves, tensions between efficiency and depth, and the sensitivity of trust to system reliability. These findings suggest that curriculum-aligned constraints and hint-first scaffolding can support instructional integration without displacing pedagogical goals. Rather than demonstrating causal learning gains, this study contributes design principles and in-situ evidence for deploying domain-specific AI assistants in technical higher-education contexts.
Or Peretz, Roei Zerahia· International Journal of Inf...· 0 citations
The proliferation of large language models (LLMs) and generative artificial intelligence has catalyzed an unprecedented pedagogical paradigm shift within higher education. This study investigates the adoption, efficacy, and cognitive impact of these technological tools among undergraduate cohorts in engineering and life sciences. Recognizing that baseline familiarity with AI does not inherently translate into advanced operational competency or prompt engineering literacy, this study evaluates the deployment of both standard LLMs and retrieval-augmented generation (RAG)-based customized assistants. The investigation employs a rigorous dual-phase methodology: an exploratory assessment of technology acceptance using standard ChatGPT, followed by a tightly controlled quasi-experiment evaluating the impact of domain-specific “GPT Custom” mentors on academic performance on complex engineering tasks. The empirical results demonstrate that customized AI assistants significantly improve final academic outcomes, yielding average grade increases of more than 15% on data-intensive analytical assignments. Furthermore, the deployment of customized assistants notably reduced grade variability among students, indicating a homogenization of academic performance that effectively levels the learning environment without compromising rigor. This improvement was statistically amplified when students utilized premium, high-capacity versions of the models for extensive synthesis tasks. Ultimately, the data indicate that while AI offers robust adaptive scaffolding, its efficacy depends on users’ critical-thinking capacities, underscoring the urgent need for educational frameworks that cultivate prompt-engineering literacy and responsible human-AI collaboration.
Jorge Cruz-Ángeles, Mariana E. Elizondo-García, Genaro Zavala et al.· International Journal of Eng...· 0 citations
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Benjamin Barlog, Hudson Craig, Zedong Peng· 0 citations
Generative AI tutors have become a common tool for independent learning, yet their capacity to support self-regulated learning (SRL) is poorly understood. This simulation-based textual analysis of prompt design evaluates a frontier large language model (Claude Sonnet 4.6) as a tutor across 60 scripted sessions on a single topic (density), crossing three levels of SRL-informed system prompting (Minimal, Moderate, Extensive) with four learner-behavior variants (Standard, Misconception, Disengagement, Overconfidence). Tutoring transcripts were scored on a 14-dimension framework spanning SRL phases, SRL developmental stages, self-determination theory principles, and Merrill’s First Principles of Instruction, applied via an LLM judge. Adding SRL context to the system prompt raised total tutoring scores, but only at the Extensive SRL support level. Minimal and Moderate prompting produced the same performance, near 36 on a 70-point scale, and Extensive prompting raised it to 40, a statistically significant effect (partial η2 = 0.24). The learner’s behavior in the session had a larger effect than the prompt did (partial η2 = 0.37), with disengaged learners scoring lowest. The threshold pattern held under an independent judge from a different developer than the tutor model. The findings support a method for evaluating GenAI tutors empirically and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
This article presents a novel approach to Intelligent Tutoring Systems (ITS) by integrating Retrieval-Augmented Generation (RAG) with Large Language Models (LLMs) to enable dynamic personalization in educational contexts. The system addresses limitations in traditional ITS that rely on static, rule-based approaches by implementing a three-layered architecture combining semantic retrieval mechanisms with generative AI capabilities. Using GPT-4 as the core LLM enhanced with a custom RAG framework, the system demonstrates improvements in response accuracy (93%), inference speed (2.1 seconds per prompt), and computational efficiency compared to a standard GPT-4 baseline, a traditional rule-based ITS, and an LLM with keyword-based retrieval. The research employs both ASSISTments (fine-grained interaction data) and EdNet (large-scale longitudinal data) datasets for evaluation. Results show that the RAG-enhanced system achieves 40% better contextual relevance compared to standard LLM implementations. The framework incorporates adaptive prompting strategies, real-time knowledge base updates, and multi-level personalization algorithms to create a dynamic educational environment.
Kuyoro Afolashade, N. Uchenna, Akinwunmi Damilare· British journal of computer,...· 0 citations