Automating Situational Judgment Tests From Likert Scales: A Multiagent LLM Framework for Literacy Assessment
Abstract
Situational judgment tests (SJTs) are a well-established class of assessment instruments that embed measurement within realistic scenarios, eliciting context-driven responses that are less susceptible to social desirability bias and more effective at capturing applied judgment than traditional self-report scales. Despite these advantages, the widespread adoption of SJTs in educational assessment remains constrained by the high cost and labor intensity of manual item development. To address this scalability bottleneck, we propose Agents for Educational SJTs (AES), a multiagent framework powered by large language models (LLMs) that automatically transforms existing Likert-scale items into psychometrically informed, context-rich SJTs. AES comprises three cooperative modules, generator, simulator, and calibrator, that emulate the expert-driven development pipeline by combining qualitative review from simulated domain experts with quantitative analysis of synthetic student responses, enabling iterative item refinement without human intervention. We instantiate AES in the domain of AI literacy, a complex and affectively rich competency that typifies the constructs for which SJTs are most beneficial. Evaluation on both human participant and LLM-simulated response datasets shows that AES-generated SJT items provide improved psychometric evidence relative to their Likert counterparts, while learners report greater engagement and perceived authenticity. These findings provide initial evidence that AES can serve as a scalable and psychometrically informed approach for supporting the development of scenario-based AI literacy assessments.