Conference
Open access
2026
Experience-driven Multi-turn Reinforcement Learning for GUI Agents
EMPO achieves substantial gains over the base model and achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o, demonstrating better generalization than prior single-turn RL approaches.
Zhengxi Lu, Jiabo Ye, Fei Tang et al.
· Annual Meeting of the Associ... · 0 citations