Skip to content
Open access

ConsensusGrade: A Human-Variability-Aware Framework for Evaluating LLM-Based Automated Grading

Aug 2026 · Applied System Innovation · 0 citations · 34 references

Abstract

Background: Evaluation of LLM-based automated grading often relies on comparison with a single human score, which can obscure meaningful variability among raters of open-ended answers. This study introduces ConsensusGrade, a consensus-aware framework that treats the human reference as a scoring envelope rather than as a single point. Methods: We analyzed 1000 open-ended student answers from 100 students across 10 questions, each graded by four evaluators. Six previously generated and aligned automated grading configurations from GradeAgentOps were compared with the four-rater human reference. The score sets were generated using Llama 3.3 70B Instruct as the primary grader, with Qwen 2.5 14B Instruct for semantic repair. Results: Human evaluators showed meaningful agreement, with ICC(A,1) = 0.712, but exact four-rater agreement occurred in only 2.2% of records. Broad score dispersion occurred in 59.0%. All automated configurations showed negative bias relative to the human median. FULL achieved 68.5% inside-envelope positioning and a chance-adjusted score of 0.454; under the central-trimmed envelope, this rate decreased to 34.3%, while configuration ordering was preserved. Conclusions: ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.

Read PDF