Self-Meta-Evolve is proposed, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successfu...
Hong-Liang Li, Lu Wang, Yong Xu et al.· 0 citations
This work introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows, and proposes a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment to enable comprehensive evaluation.
Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.
Hongliang Li, Yijin Liu, Zhiwei Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.