Can Generative Large Language Models Serve as Raters for Test Development? A Systematic Evaluation Across Tasks, Models, and Inference Configurations
The present study investigates the effectiveness of generative large language models (LLMs) as raters across three common rating tasks: (a) social desirability ratings, (b) content validity ratings, and (c) trait importance ratings. Specifically, we examine reliability and validity of LLM-generated ratings across varying occupational contexts, rating methods, LLM families (i.e., GPT-4, GPT-5, and Sonnet 4.5), and inference configurations (i.e., prompt design and temperature settings). Results indicate that LLM ratings exhibit strong reliability and convergent validity in social desirability ratings across occupational contexts, as well as acceptable convergence with human ratings in Likert-type content validity evaluations. In contrast, reliability and convergent validity for trait importance ratings were inconsistent across occupational contexts. Variations in prompt design and temperature settings generally produced small to negligible effects on reliability and validity. Overall, the findings suggest that LLMs can function as effective supplementary raters in test development and validation processes, although greater caution is warranted for certain rating tasks. Practical implications and directions for future research are discussed.