How Prompt Design Shapes AI-Assisted Assessment: Reliability, Validity, and Learning Implications
Abstract
This study examines the reliability and validity of generative AI in summative assessment, emphasizing how prompt design influences grading when applying a common rubric to complex student work. Fifteen business plans from a master’s-level course were evaluated by GPT-5 through Microsoft Copilot using three prompts of increasing rigor (basic, intermediate, rigorous). Each plan was scored in five independent runs per prompt, producing 225 AI evaluations. Analyses included intraclass correlation for consistency, severity contrasts, and convergence with instructor scores using correlation, error metrics, and Bland–Altman limits of agreement. Prompt design significantly shaped score distribution and strictness. The most rigorous prompt reduced inflated scores and aligned more closely with instructor judgments, yet it also underestimated performance and showed the greatest inconsistency. Single AI runs were unreliable, but averaging multiple evaluations improved stability. At the criterion level, AI struggled to match instructor ratings on commercial and economic viability, even under stricter prompts. Findings highlight the pedagogical implications of prompt sensitivity in AI-assisted grading. Reliability and fairness depend not only on rubric quality but also on evaluative instructions. Results support multi-run aggregation, bias-aware calibration, and hybrid human–AI models to ensure rigor and equity in technology-enhanced assessment. These findings inform the design of AI-enhanced assessment practices that support fair, transparent, and pedagogically aligned learning environments.