GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI acti...
Zi-Qiang Wang, Li Gu, Zhixiang Chi et al.· 0 citations
This work introduces MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories, and reveals benchmark quality metrics reliably predict real-world judge utility.
Zi-Qiang Wang, Li Gu, Zhixiang Chi et al.· 0 citations
A new FSTT-DA framework that integrates LoRA fine-tuning with model merging and proposes a hypernetwork trained via meta-learning that generates per-column merging factors to combine LoRA modules to adapt the learned knowledge to a specific target domain.
Siobhan Reid, Zhixiang Chi, Li Gu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.