Skip to content
Open access

Algorithm Aversion Towards AI-Generated Multiple-Choice Questions

Matthias Carl Laupichler Susanne Eigster Seifollah Ahmadi Ilona Grunwald Kadow Tobias Raupach Johanna Flora Rother
Aug 2026 · The journal of the International Association of Medical Science Educators : JIAMSE · 0 citations · 33 references

Abstract

Algorithm aversion and anti-AI bias describe the human tendency to evaluate decisions made by human agents more favorably than those made by artificial intelligence (AI) algorithms. This study aimed to examine whether algorithm aversion influences students’ evaluations of multiple-choice questions (MCQs) as well as the psychometric quality of items generated by a large language model (LLM). Additionally, the authors explored the possible influence of AI literacy and attitudes toward AI. The research team administered a formative examination consisting of two blocks. Block 1 was labeled “authored by a medical educator,” although 50% of the 20 items were in fact generated by an LLM. Block 2 was labeled “generated by an LLM,” but likewise included 50% human-authored items. The human-authored and AI-generated questions were based on the same 20 learning objectives. Following each item, medical students rated its perceived relevance for exam preparation and its quality of wording. AI literacy and attitudes toward AI were assessed using validated instruments. Participants solved significantly more LLM-generated items ( M = 77%) correctly than human-authored items ( M = 48%), U = 34, p < .001, r = 0.71. No significant differences were observed in discriminatory power. Cumulative link mixed models revealed that students rated the relevance of actually AI-generated items more positively. In addition, they exhibited an anti-AI bias in their ratings of item wording quality, as items labeled as LLM-generated were rated significantly worse, irrespective of their actual origin. Although the overall findings statistically suggest an anti-AI bias, their practical implications remain unclear, as relevance and wording quality were rated very positively overall.

Read PDF