We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI acti...
Zi-Qiang Wang, Li Gu, Zhixiang Chi et al.· 0 citations
StepReflect is proposed, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence, and established as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
Lin-Qiang Guo, Wei Liu, Li Gu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.