Skip to content
Open access

How consistent is the algorithm? Examining the intra- and inter-rater reliability of LLM-based writing assessment

Aug 2026 · Frontiers in Education · 0 citations · 27 references

Abstract

Unlike constrained constructed-response assessments, the reliability of essay-based assessments has long been debated in language education research. Recently, AI—particularly large language models—has been proposed as a tool to enhance scoring reliability. This study examines the reliability of LLM-based scoring in writing assessment. OpenAI's ChatGPT assessed 192 essays written by EFL learners using a criterion-based analytic rubric covering task achievement, grammatical range and accuracy, lexical resources, organization, and mechanics. Scoring was replicated after three weeks under similar conditions. The two rounds were compared to assess intra-rater reliability and then contrasted with human ratings to assess inter-rater reliability. Results indicated high intra-rater reliability, particularly for task achievement and organization. Lower consistency was observed for mechanics and grammatical accuracy, suggesting a need for clearer criteria or further model training. Human–AI comparisons showed slight to fair exact agreement (Cohen's kappa) and moderate agreement when accounting for the degree of disagreement.

Read PDF