Skip to content

Author

Iván Barcia-Santos

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Semantic anchoring with concise ideal answers outperforms unstructured full course materials as context for multi-LLM automated grading of open-ended questions

Abstract Large language models (LLMs) are increasingly used to grade open-ended student responses, yet the role of contextual input in this process remains poorly understood. This study compares three context conditions for multi-LLM automated grading: no context, full course materials, and instructor-defined ideal answers as semantic anchors. Using a dataset of 3,041 student responses (3,011 after common-support exclusions) to 50 open-ended questions from an undergraduate computer science course, we evaluated three base LLMs (DeepSeek, Qwen, Gemini) against grades derived from two independent blind instructor assessments. A factorial analysis based on the Aligned Rank Transform revealed significant main effects of model and condition, with a significant interaction. Ideal-answer anchoring significantly outperformed both alternatives in absolute grading error, while providing full course materials significantly worsened accuracy relative to the no-context baseline. The anchored condition achieved the lowest mean absolute error (1.268), the highest correlation with instructor grades ( r = 0.801), and the lowest inter-model disagreement (median SD = 0.864), at a per-response cost comparable to the no-context baseline (EUR 0.00119 vs. 0.00113) and 4.2 times cheaper than the full-materials condition (EUR 0.00501). The absolute accuracy gain over the no-context baseline is small (≈ 0.08 points on a 0–10 scale; marginal R 2 = 0.013); its practical value lies in the convergence of accuracy, inter-model agreement, feedback-quality and cost improvements and in avoiding the accuracy loss caused by unstructured full course materials. A complementary analysis of 27,099 feedback instances, validated against a human gold standard (κ = 0.898), showed that out-of-scope feedback decreased by 25.4% under semantic anchoring, with a logistic regression revealing that this benefit was concentrated in two of the three evaluators. Evidence derives from a single course, institution, language (Spanish) and academic year, and from convergent-answer theoretical assessment; within this setting, concise instructor-defined ideal answers (rather than large volumes of unfiltered course material) yielded the most reliable grading and feedback.

Jorge Cisneros-González, Natalia Gordo-Herrera, Iván Barcia-Santos et al. · 0 citations