LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards
Abstract
Large classes in Chinese College English programmes make frequent analytic assessment of student writing difficult. Large language models (LLMs) may support more frequent formative assessment, but their scores may vary across queries and be systematically harsher or more lenient than local teacher ratings. Using a corpus-based, five-fold cross-validated comparative rater-evaluation design, this study examined whether statistical calibration could make LLM-assisted scores more interpretable for College English writing assessment and where their use should remain limited. Data comprised 414 timed argumentative essays written by Chinese non-English majors at one applied undergraduate institution. Two trained College English teachers independently rated the essays on a seven-dimension analytic rubric informed by China’s Standards of English Language Ability, providing the local reference standard. Three LLMs rated each essay–dimension pair on five occasions. Under five-fold cross-validation, uncalibrated scores were compared with location–scale correction, isotonic calibration, and equipercentile linking, using quadratic weighted kappa, Spearman correlation, mean absolute error, signed bias, and half-point tolerance accuracy. Agreement between models did not imply agreement with teachers: two models showed inter-model kappa values of 0.70–0.78 but an average kappa of only 0.15 with teacher ratings while rating the essays about one band more severely. Calibration removed most of this severity difference and raised pooled kappa to 0.61–0.70 depending on the method (0.63–0.64 under equipercentile linking), compared with a teacher–teacher agreement benchmark of 0.747. The three methods differed little, and the improvement mainly reflected closer alignment of score distributions rather than better judgement of writing quality. Agreement was higher for vocabulary, syntax, and grammar but remained low for cohesion and conventions. The findings suggest that LLM-assisted scoring may support low-stakes formative feedback when calibrated to local teacher standards and used under teacher supervision, while teachers retain responsibility for judging content, coherence, argumentation, and communicative quality.