See, a two-stage data synthesis framework consisting of an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling, yields reproducible and explainable data generation.
Abstract
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. Existing pipelines often rely on costly human demonstrations or on-policy framework, which tends to over-sample common flows while missing rare transitions and complex multi-step procedures. To address this problem, we propose SEE, a two-stage data synthesis framework consisting of (i) an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and (ii) a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling. This design yields reproducible and explainable data generation, while explicitly preventing spurious cycles and enabling long-horizon composition. Across multiple real-world apps, SEE produces trajectories with an average length of 14.8 steps while avoiding spurious loops, and agents fine-tuned on SEE achieve improved task success and generalization to unseen screens. We will publicly release our synthesis code and dataset.
StepReflect is proposed, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence, and established as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
EMPO achieves substantial gains over the base model and achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o, demonstrating better generalization than prior single-turn RL approaches.
Zhengxi Lu, Jiabo Ye, Fei Tang et al.· Annual Meeting of the Associ...· 0 citations
This work presents Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts and shows that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized.
M. Venkatesan, Shicheng Wen, Jiajing Guo et al.· 1 citation
LocalLSTC is introduced, a training-free architecture that organizes control by temporal scope, maintaining persistent cross-step state to guide short-term execution commitments, and identifies temporal organization of control information as a distinct architectural dimension for locally deployed GUI agents.
Qwen-UI-Agent is presented, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments, that sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models.
Hanzhang Zhou, Panrong Tong, Xu Zhang et al.· 2 citations
This work proposes a closed-loop Multi-Agent Simulation Framework to synthe-size diverse, faithful, and policy-aligned shopping trajectories, and presents synthetic data that enables a small model to significantly outperform same-size baselines and surpass a large-model baseline.
Qing Ping, Changyou Chen, Binxuan Huang· Proceedings of the 64th Annu...· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026