Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
GMA is presented, a benchmark for evaluating general mobile assistants in challenging real-world scenarios, and shows that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models.
Yiqi Zhu, Feiyu Gao, Jiakang Fan et al.
· 0 citations