This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf, enabling a systematic study of models's strategic planning, adaptation, and robustness under uncertainty.
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu et al.· 0 citations
Experimental results on a broad range of MLE tasks with diverse model types and scales demonstrate that Matryoshka Agent is an effective and scalable paradigm for long-horizon MLE tasks and complex agentic problem solving.
Rushi Qiang, Changhao Li, Haotian Sun et al.· arXiv.org· 0 citations
A controlled study of large language model agents across 260 configurations shows when multi-agent collaboration helps or hurts performance, and introduces a predictive model that selects the best architecture in 87% of held-out within-domain configurations.
Y. Kim, Ken Gu, Chanwoo Park et al.· Nature Machine Intelligence· 6 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.