Skip to content
Preprint

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

Aug 2026 · 1 citation · 50 references
Computer Science

TL;DR

Social Gym is introduced, an environment of 21 multi-agent social games whose rule-decided outcomes make agent performance verifiable and objective, and SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop that offers a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.

Abstract

LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.

View source

Similar papers

#small language model Book Open access Sep 2026

Evaluating LLM Social Cognition Through Multi-Agentic Strategic Games

Large language models (LLMs) now power the reasoning core of intelligent virtual agents deployed across an expanding range of social settings, from tutoring students and supporting patients in healthcare, to mediating group discussions and representing humans in various social settings. Effective deployment demands soc...

Kevin Kurian, Kevin Scroggins, Emmanuel Dorley et al. · 0 citations
Book Open access Sep 2026

Games of Diplomacy: Eliciting Spite Against Automated Opponents Through Structured Game Environments

Social games require a nuanced understanding of human behavior and are uniquely difficult to solve computationally, even with access to large behavioral datasets. Examining specific aspects of social gameplay, however, can simultaneously yield insight into human behavior and inform the design of more accurate, human-li...

Noah Ari, Darryl Roman, Nusrath Jahan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs

Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent frameworks often have every model work toward a...

Kunal Jha, Max Kleiman-Weiner, Natasha Jaques · 0 citations
#artificial intelligence Preprint Sep 2026

GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games

Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluate...

Dhananjay Ashok, Adam Shen, A. Feng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games

Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly...

Yechan Hwang, S. Bae, Jeongmo Kim et al. · 0 citations
2025

Collaborative Reasoner: Self-Improving Social Agents with Synthetic Conversations

This work presents Collaborative Reasoner, a framework to evaluate and improve the collaborative reasoning abilities of language models, and proposes a self-play method to generate synthetic multi-turn preference data and further train the language models to be better collaborators.

Ansong Ni, Ruta Desai, Yang Li et al. · 7 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.