Do AI Grading Systems Systematically Differ from Human Teachers’ Grading? Evidence of Bias and Consistency in Educational Assessment
The rapid development of artificial intelligence (AI) has introduced new possibilities for transforming educational assessment processes. Among these developments, AI-assisted grading systems have attracted increasing attention due to their potential to improve efficiency, consistency, and scalability of student evaluation. The present study examines the role of artificial intelligence in student grading by comparing AI-generated scores with human teacher evaluations and by exploring teachers’ perceptions regarding the use of AI in educational assessment. The research adopts a quantitative comparative design. Student-written responses were independently evaluated by teachers and AI systems, and the resulting scores were statistically analyzed to examine the level of agreement between the two grading approaches. In addition, a structured questionnaire was administered to teachers to investigate their attitudes toward AI-assisted grading. The findings indicate that while some AI systems produce scores comparable to human evaluators, others exhibit statistically significant differences, highlighting variability across models. Furthermore, AI systems were found to produce more consistent grading outcomes in relation to the corresponding human evaluators. Nevertheless, teachers recognized the potential of AI to reduce the time required for assessment tasks. However, concerns related to fairness, transparency, and the interpretation of complex student responses remain important considerations. Overall, the results suggest that artificial intelligence can effectively support educational assessment when implemented within hybrid evaluation models that combine automated analysis with human pedagogical oversight.