Structured Human-AI Teaming for UX Heuristic Evaluation with Human-in-the-Loop Supervision
Abstract
UX heuristic evaluation is a core human factors method for assessing interface designs, but expert-led approaches are constrained by expert availability and time demands. Although multimodal large language models can automate heuristic evaluation from screenshots, full automation raises concerns about inaccurate design interpretation, unreliable reasoning, and limited transparency. This study proposes a structured human-in-the-loop architecture that reframes heuristic evaluation as a problem of cognitive labor distribution. Two AI agents—the Design Representation Generator and the Heuristic Evaluator—collaborate with a human supervisor. The supervisor reviews and corrects intermediate design representations and challenges, contextualizes, or refines AI-generated UX issues through correction and adjustment loops. Across four Android mobile app task scenarios, the human-in-the-loop system was compared with a fully automated baseline using the same two-agent pipeline without human supervision. The human-in-the-loop system produced heuristic evaluation results that showed substantially greater alignment with human experts than those generated by the automated baseline.