Skip to content
Open access

Evaluating large language models for rubric-based essay grading in an undergraduate biology course

Jul 2026 · Journal of Microbiology & Biology Education · Vol 27 · 0 citations · 14 references
Medicine

TL;DR

It is suggested that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components, and this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.

Abstract

ABSTRACT This study examines how three large language models (LLMs), ChatGPT, Claude, and Gemini, assign grades to undergraduate-level essays in a biology course using a standardized rubric. Each LLM evaluated a data set of 200 essays under two prompting conditions: zero-shot (uncalibrated) and few-shot (calibrated using a small set of exemplar essays). LLM-assigned scores were directly compared with instructor-assigned scores, showing only moderate alignment with instructor grading, with variability observed across models and prompting strategies. Differences in grading behavior were also evident with different items on the rubric, with higher alignment for structural writing components and lower alignment for content- and reasoning-based criteria. Additionally, LLMs showed greater agreement with instructor scores than with one another, indicating substantial inter-model variability under identical grading conditions. These findings suggest that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components. In this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.

Read PDF

Similar papers

Open access Aug 2026

EduFairBench: reproducible evaluation of large language models for educational assessment

Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al. · 0 citations
#small language model Preprint Aug 2026

Grading Needs a Rubric, Not Intelligence

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.

Jhen-Ke Lin · 0 citations
Open access Aug 2026

LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards

Large classes in Chinese College English programmes make frequent analytic assessment of student writing difficult. Large language models (LLMs) may support more frequent formative assessment, but their scores may vary across queries and be systematically harsher or more lenient than local teacher ratings. Using a corpus-based, five-fold cross-validated comparative rater-evaluation design, this study examined whether statistical calibration could make LLM-assisted scores more interpretable for College English writing assessment and where their use should remain limited. Data comprised 414 timed argumentative essays written by Chinese non-English majors at one applied undergraduate institution. Two trained College English teachers independently rated the essays on a seven-dimension analytic rubric informed by China’s Standards of English Language Ability, providing the local reference standard. Three LLMs rated each essay–dimension pair on five occasions. Under five-fold cross-validation, uncalibrated scores were compared with location–scale correction, isotonic calibration, and equipercentile linking, using quadratic weighted kappa, Spearman correlation, mean absolute error, signed bias, and half-point tolerance accuracy. Agreement between models did not imply agreement with teachers: two models showed inter-model kappa values of 0.70–0.78 but an average kappa of only 0.15 with teacher ratings while rating the essays about one band more severely. Calibration removed most of this severity difference and raised pooled kappa to 0.61–0.70 depending on the method (0.63–0.64 under equipercentile linking), compared with a teacher–teacher agreement benchmark of 0.747. The three methods differed little, and the improvement mainly reflected closer alignment of score distributions rather than better judgement of writing quality. Agreement was higher for vocabulary, syntax, and grammar but remained low for cohesion and conventions. The findings suggest that LLM-assisted scoring may support low-stakes formative feedback when calibrated to local teacher standards and used under teacher supervision, while teachers retain responsibility for judging content, coherence, argumentation, and communicative quality.

Yong-Ping Wang, Ning Liu, Xi-Zhi Chu et al. · 0 citations
Preprint Jul 2026

When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring

Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target essays are unseen during training. We further find that increasing target-essay supervision improves performance, with our best fine-tuned open-source Llama-based model outperforming GPT-5-mini prompting by 2.1% macro F1 and trailing GPT-5 by 1.9%. These results show that trait-based intermediate structure and controlled supervision improve generalization to unseen rubrics.

Nischal Ashok Kumar, Payu Wittawatolarn, Sana Kang et al. · 0 citations
Preprint Jul 2026

Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in"AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models"(Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.

John Maurice Gayed · 0 citations
Open access Jul 2026

Effect of large language model assistance on undergraduate art history question-answering performance: a randomized crossover pilot study

Introduction Large language models (LLMs) are increasingly used in higher education, yet empirical evidence for their effectiveness in art education remains scarce. This study aimed to evaluate whether LLM assistance could improve undergraduate art history question-answering performance and explanatory support. Methods This study developed the Art History Theory Question Set (AHTQS), comprising 104 single-choice items with Bloom-level annotations, and benchmarked three LLMs (ChatGPT-4o, DeepSeek-V3, and Qwen2.5-Plus). DeepSeek-V3 showed the highest accuracy (96.2%) and lowest observed run-to-run variability and was selected for a randomized crossover pilot study with six undergraduates. The primary outcome was the change in examination accuracy from independent to LLM-assisted answering. A Likert-scale evaluation involving nine students and three instructors was also conducted to assess the clarity and coherence of LLM-generated explanations. Results A one-sided Wilcoxon signed-rank test showed significant improvement with LLM support [W(6) = 21.0, p = 0.0156, r = 0.879], with median accuracy increasing from 43.3% to 93.3% (median gain = 42.3%). Five of six students showed higher accuracy under the LLM-assisted condition, and no clear evidence of a sequence or carryover effect was detected (Mann-Whitney U = 7.0, p = 0.3758). Domain-level analyses indicated significant gains in all four categories (p < 0.05). Error-frequency analysis further showed marked reductions in high-frequency mistakes. The Likert-scale evaluation indicated high perceived clarity and coherence of LLM explanations, with favorable but more cautious instructor ratings. Discussion These pilot findings suggest that supervised LLM assistance may support art history question-answering and explanatory feedback. Future studies should validate these findings in larger cohorts, assess delayed learning retention, and examine open-ended, image-based, and higher-order art history tasks before curriculum-level implementation.

Yunting Zhang, Fan Zhang, Zili Zhang · 0 citations