Social Gym is introduced, an environment of 21 multi-agent social games whose rule-decided outcomes make agent performance verifiable and objective, and SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop that offers a reproducible, verifiable foundation for measuring and improving LLM socia...
Rapid development of Large Language Models (LLMs) and similar automated approaches for translation tasks is increasingly affecting the landscape of translation technologies. As concerns about the outsourcing of translator work to these automated translation tools grow, it is increasingly crucial to gather insights from...
Daniel Chechelnitsky, Sireesh Gururaja, Seyi Olojo et al.· 0 citations
When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g.,"How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception ("regex are fragile") and meaningfully direct the user toward a pragmatic solution that will addres...
Akhila Yerukola, Jena D. Hwang, Ming-Qian Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.