Skip to content

KSA Profiles for Standard Setting: A Generative AI Approach to Empirically Grounded Cut Score Evaluation

Aug 2026 · Journal of Educational Measurement · Vol 63 · 0 citations · 20 references

TL;DR

Generative AI is applied to AP U.S. History and AP World History essay data to demonstrate how this kind of evidence could help panels evaluate whether a proposed cut score captures the distinctions a policy intends and inform discussion of proposed cut scores or performance expectations considering evidence of the KSAs students demonstrate.

Abstract

Standard setting for essay‐based assessments often relies on expert judgment and performance‐level descriptors, but these sources do not always show how examinees at adjacent score levels differ in the knowledge, skills, and abilities (KSAs) they demonstrate. This study introduces a method that uses generative AI to fill that gap. A large language model reads student essays and estimates the likelihood that each student demonstrated a set of pre‐defined knowledge, skills, and abilities (KSAs). Those estimates are then summarized as KSA profiles—representations of which competencies are characteristically more or less evident within a score region or at a proposed cut score boundary. To demonstrate the method, we apply it to AP U.S. History and AP World History essay data using hypothetical cut scores as an exploratory example. Results show that profiles differed across score levels, that the same broad KSA structure emerged across both subjects, and that AI‐derived measures aligned with established indicators of historical reasoning. We discuss how this kind of evidence could help panels evaluate whether a proposed cut score captures the distinctions a policy intends and inform discussion of proposed cut scores or performance expectations considering evidence of the KSAs students demonstrate.

View source

Similar papers

Preprint Aug 2026

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs'underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.

Alona Strugatski, Licol Zeinfeld, Giora Alexandron · 1 citation
Preprint Jul 2026

Generative AI Availability, Grades, and Student Satisfaction at a Large University

It is found that there is no significant differential effect of GenAI availability on grades overall or among previously lower-performing students, and the findings temper concerns that GenAI inflates grades and reduces students's satisfaction.

J. Dumlao, Meng Wang, Zhonghan Xie et al. · 0 citations
Open access Aug 2026

RoCulturaMCQ: Building a Benchmark While Learning Statistics

The broad adoption of Large Language Models (LLMs) has increased the need for human-curated datasets that serve as evaluation benchmarks. This need is particularly pronounced for non-English languages and for tasks that are inherently subjective and require multiple human perspectives. One such example is the development of benchmarks designed to assess the cultural awareness of LLMs. Statistics and data science courses offer a potential setting for developing such benchmarks while teaching students to apply LLM evaluation techniques using statistical inference. This paper presents a pilot project in which students in a statistics course within a data science engineering program created culturally diverse multiple-choice questions, generated answers using LLMs, and applied statistical methods to assess model accuracy. Student feedback indicated the project was engaging and useful for learning, while also highlighting a notable reliance on LLMs, particularly for interpreting statistical results. The resulting dataset comprises 1355 multiple-choice questions across 18 categories, including language, social media, and politics. After filtering valid items, the dataset was used to evaluate both closed- and open-source LLMs. Results show that the Gemini (closed-source) and Qwen (open-source) model families achieved the best performance, with improvements linked to model size, reasoning capabilities, and access to search tools. The best closed-source model achieved an accuracy of 97.66%, whereas the best open-source model achieved an accuracy of 79.07%. Qualitative analyses of errors in the filtering procedure and model reasoning process point to possible explanations into the challenges LLMs face when handling culturally specific content. Furthermore, results support a cultural injection hypothesis, whereby cultural knowledge is embedded during pretraining and accessed through instruction tuning. Through this work, we aim to demonstrate how statistics and data science courses can provide productive contexts for developing open-source benchmarks for non-English languages while also enriching students’ learning experiences. The dataset is publicly available.

Denis Iorga, Razvan Muntean, Mihai Masala et al. · 0 citations
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
Open access Jul 2026

Evaluating large language models for rubric-based essay grading in an undergraduate biology course

It is suggested that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components, and this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.

M. Naidu, Nikolas S. Montaquila, Jessica P Roa et al. · 0 citations
Review Aug 2026

Evaluation in the Age of AI: Output as Evidence of Learning

The rapid adoption of artificial intelligence (AI), particularly large language models (LLMs), has fundamentally disrupted how learning is demonstrated and evaluated in higher education. Tasks that once served as proxies for understanding-such as writing essays, solving problem sets, or producing computer code-can now be generated superficially by AI systems with minimal human effort. This paradigm shift raises a critical ethical question: how should learning be evaluated when traditional indicators of competence are easily outsourced? This paper examines the ethical challenges of educational evaluation in the age of AI from a university-level perspective. We argue that the core problem extends beyond academic dishonesty to a deeper misalignment between assessment practices and the learning outcomes they are intended to measure. Evaluation regimes that rely on artificial constraints risk measuring compliance, access, or concealment rather than genuine understanding, reasoning, or judgment. By analyzing institutional responses and presenting empirical survey data, we highlight the need for alternative assessment models that emphasize process over product. The goal is to establish ethically informed assessment strategies that preserve student agency and accountability in an automated age.

Md Zarzees Uddin Shah Chowdhury, Samin Khan · 0 citations

Related blog posts