Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation
Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at once, leading to substantial context overload. Moreover, they primarily focus on task comp...